Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety


Abstract

Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child sexual abuse material, facilitate child sexual exploitation, and reduce barriers to harm. In this paper, we argue that protecting children from AI-facilitated sexual abuse requires new approaches to AI safety. Existing safety techniques assume data accessibility, transparency, and evaluation practices that are incompatible with the ethical and legal constraints surrounding child sexual abuse material. We examine how these constraints create new technical challenges, such as limitations on dataset auditing, red teaming, and fine-tuning prevention. In turn, we outline 15 open problems in online child sexual exploitation and abuse across the AI development lifecycle, from dataset curation and model design to deployment and long-term maintenance. We propose targeted recommendations for researchers, developers, and policymakers to bridge the gap between theoretical AI safety and the realities of child protection. Our work aims to reframe preventing AI-facilitated child sexual abuse as a central, safety-critical dimension for AI research, motivating work that translates responsible AI principles into concrete safeguards against the exploitation of children.

1 Introduction↩︎

Artificial intelligence (AI) systems have achieved remarkable capabilities in content creation. However, these same capabilities increasingly pose risks to children, including facilitating child sexual abuse and exploitation (CSAE), particularly via the development of AI-generated child sexual abuse material (AIG-CSAM). This misuse represents a critical challenge at the intersection of AI safety, child protection, and technical system design—one that most existing AI safety frameworks inadequately address [1][4].

The scope and severity of this problem are substantial. The National Center for Missing and Exploited Children (NCMEC) received more than 440,000 reports of AI-generated material related to CSAE in the first half of 2025 alone [5], compared to 7,000 reports in 2023-2024, and the Internet Watch Foundation (IWF) observed a 400% increase in AIG-CSAM reports since 2024 [2], [6]. A recent study found that 13% of U.S. sexual extortion victims reported that the perpetrator used AI to create blackmail material [7], and a 2025 survey of Australian young people found that 41% of sexual extortion victims reported that the material used to blackmail them had been digitally manipulated [8]. In the U.S., 1 in 17 teenagers aged 13–17 report having been victimized by deepfake nude images [9]. AIG-CSAM has thus been recognized by the AI safety community as a critical area for intervention, with immediate, quantifiable marginal risk [10].

However, despite the known risks of AIG-CSAM, many existing techniques from AI safety and secure & privacy-preserving machine learning cannot be directly used for AIG-CSAM prevention, detection, and mitigation. Paradoxically, in part because child safety is an important and established area of concern, there are unique constraints and restrictions surrounding CSAM access/usage. These restrictions are ethical and necessary, but create complications for researchers and practitioners seeking to develop and assess AI safety solutions. This motivates our main position (below), which we explore in detail in this work:

Main Position Existing AI safety research makes assumptions that do not match the legal and ethical constraints of safety applications preventing child sexual abuse and exploitation. Bridging this gap by solving open sociotechnical problems will be critical to make AI safety solutions effective for child safety in practice.

Table 1: CSAM access, usage, and generation capabilities for various actors. The inability to access CSAM as well as generate and train on CSAM data can limit the application of existing tools in AI safety. *While AI developers and providers generally can’t access or train on CSAM, there are specific, limited cases where access may be permissible (see Section [sec:background:access]).
General Public AI Developers/Providers Reporting Hotlines Law Enforcement
Can legally access CSAM No No* Yes Yes
Can access hashes No Yes Yes Yes
Can train on CSAM No No* Yes Yes
Can generate CSAM No No No Yes

We outline current industry consensus on solutions to address AIG-CSAM [11], pointing out gaps between assumptions made by state-of-the-art mitigations and practical limitations. We then highlight 15 open problems (9 in the body, 6 in Appendix 10) spanning model development (\(\S\)3), deployment (\(\S\)4), and maintenance (\(\S\)5), which represent critical research and policy priorities where advances could have immediate impact on protecting children. Finally, we provide recommendations for researchers, AI providers, and policymakers to address these open problems and bridge ecosystem gaps.

2 Background↩︎

Generative AI enables non-experts to create realistic digital media at unprecedented scale. The AI safety community has raised concerns about resulting risks, such as bias propagation [12], disinformation [13], bioterrorism [14], job displacement [15], and cybercrime [16]. In this work, we specifically examine generative AI risks for child safety, focusing on photo-realistic child sexual abuse material (CSAM). AIG-CSAM presents critical harms to child safety. It complicates victim identification by expanding the volume of material that law enforcement must navigate to locate children in abuse scenarios [1], [4]. It enables re-victimization as bad actors can fine-tune models on existing CSAM to generate content imitating specific victims while producing new poses and acts of violence [1], [4]. It also broadens the pool of victims for sexual extortion, with offenders using image editing tools to sexualize benign depictions of children for such schemes [17].

Both preventive and reactive measures around AIG-CSAM are necessary to improve child safety: civil society [18], industry [19], and regulatory bodies [20] have all pursued relevant efforts. However, as we show, CSAM imposes unique constraints on existing AI safety techniques. To illustrate the resulting open research gaps, in this section we first describe what makes AI child safety particularly challenging (\(\S\)2.1) and detail key organizations involved in intervention efforts (\(\S\)2.2). Disclaimer: Throughout this work, we note that we focus on AI-generated imagery as the key modality of interest, and primarily take a U.S./western perspective when considering legal constraints that may impact safety techniques. We discuss limitations, broader perspectives, and alternate views in \(\S\) 6 and \(\S\) 7, and provide a detailed discussion of related work in App. 9.

2.1 What makes AI child safety challenging?↩︎

The illegal and sensitive nature of CSAM, alongside other guardrails related to children’s data, creates unique challenges for AI safety. Below we highlight several underlying challenges which we will revisit throughout the paper:

Data restrictions. CSAM is illegal in most countries, with legislation such as COPPA and GDPR establishing strict privacy safeguards for children’s data [21][23]. As we show, these data access restrictions are at odds with many standard AI safety approaches, which require access to training and/or evaluation data.

Evaluation restrictions. Evaluation and red teaming methods typically rely on model prompting to assess model behavior and capabilities, but intentionally generating AIG-CSAM is also illegal in the U.S.  [24]. Similarly, standardized CSAM evaluation datasets do not exist, as it is illegal to access this data. — Adversarial and opaque environment with strict guarantees. Child safety demands strong guarantees, as AIG-CSAM output is intrinsically illegal in most countries and can directly harm specific children. However, these guarantees can be difficult to achieve, as CSAE offenders collaborate to actively circumvent safeguards, while AI system developers typically only disclose high-level descriptions of safety mechanisms (\(\S\)10.2).

Wellness implications. CSAM exposure and the psychological demands of adversarial assessment pose substantial risks, including trauma and PTSD [25]. This can make it challenging to apply safety techniques requiring significant human involvement.

2.2 Organizations and access assumptions↩︎

Three key groups make up the AIG-CSAM prevention and response ecosystem (see Table 1 for a summary).

AI Developers and Providers. AI developers (individuals/organizations that build AI technology) and AI providers (platforms that host AI tools) both have CSAM reporting and retention obligations under U.S. law [26]. They can access CSAM hashes via institutions like NCMEC for matching and reporting. In certain cases, they may be allowed to extract embeddings from CSAM detected on their platforms to train detection methods [27]*, but direct access would require partnership with law enforcement or reporting hotlines.

Reporting Hotlines. Hotlines (e.g. NCMEC, IWF) are authorized to receive and process CSAM reports, to facilitate removal from online platforms and support LE investigation and prosecution efforts. These organizations legally house CSAM, enabling hash sharing, victim ID, and abuse prevention efforts. While they can provide scoped, secure access to CSAM to other institutions building CSAM prevention and detection technology, they cannot directly assess generative AI models for CSAM capabilities.

Law Enforcement (LE). Finally, LE has the legal authority to investigate, collect, and analyze CSAM as evidence while working to identify victims, apprehend perpetrators, and disrupt child exploitation networks. However, although they have the ability to train detection methods and assess AI models for CSAM capabilities, they may lack the budget or technical expertise for these efforts.

As we will see, the restrictions surrounding CSAM access for these various groups (particularly for AI developers/providers) have wide-ranging impacts on the use of existing safety tools related to AI system development (\(\S\)3), deployment (\(\S\)4), and maintenance (\(\S\)5).

3 Developing Safe Models by Design↩︎

Developing safe generative AI models requires proactively addressing child safety risks before and during training. Large-scale datasets used to train generative models are often scraped from the Internet with minimal cleaning [28], leading to cases where datasets used to train popular models were found to contain CSAM [29], [30]. Generative models may also produce harmful outputs through concept fusion, combining attributes from separate training examples (e.g., adult pornographic content and benign child depictions) [31], [32]. Furthermore, by using generative models to partially edit CSAM, offenders can make it difficult to determine whether material depicts active abuse, recovered victim-survivors, or deepfakes.

3.1 Current Approaches↩︎

Currently, model developers can use allowed/disallowed website lists for data curation to avoid sites hosting CSAM, [33], and can detect CSAM in training data by hash matching against third-party CSAM databases and using CSAM classifiers [18], [27], [34]. One approach to address concept fusion is NSFW filtering, which considers removing adult material from datasets prior to model development [11].

Developers also use red teaming to find harmful content queries, patching exploits before model release [35], [36]. Directly prompting for AIG-CSAM is illegal in the U.S., so developers report either testing proxy concepts [11] or avoiding such testing entirely [37].

Content provenance solutions have also been implemented to support victim ID efforts. These tools provide a feedback mechanism for researchers, AI providers, and policy makers. By allowing stakeholders to become aware of the models that are producing problematic output, systemic issues can be addressed in existing mitigations. Current approaches include C2PA [38] (an open metadata-based standard for digital content origin and edits) and watermarking [39][41].

3.2 Open Problems↩︎

Open Problem A1: Partial data cleaning How does data cleaning affect a generative model’s ability to depict a concept? To what degree can partial cleaning guarantee that a model cannot depict CSAM?

Guaranteeing complete CSAM removal is difficult due to imperfect detection technology, scale of data, and moderator wellness concerns. Determining whether imperfect removal (e.g., 99.9%) can prevent text-to-image models from learning to generate CSAM is an important open question. However, direct analysis of CSAM filtering strategies is challenging due to data access restrictions.

Existing work. LLM [42], [43] and diffusion studies [44] show that filtering training data can minimize harmful capabilities in the model, and research in text-to-image generation shows a critical number of training samples are required for concept composition [31], [45].

Limitations. AIG-CSAM generation requires strong safety guarantees. While existing work is a starting point to study the effectiveness of data cleaning, the AIG-CSAM problem would benefit from formal guarantees on a model’s ability to generate harmful content. Moreover, because it is illegal for researchers to generate CSAM, it is not yet clear how to exhaustively test models’ capabilities to ensure they are safe in these scenarios.

Open Problem A2: Preventing concept fusion Can generative models be selectively blocked from combining high-risk concepts, such as children and NSFW material?

A unique concern for CSAM is that harmful content can also be created by composing two potentially benign concepts (e.g., benign child depictions and adult content) (see example experiments in \(\S\)3.2.1). Addressing this via comprehensive data filtering is challenging for similar reasons to CSAM removal. In addition understanding whether partial data cleaning prevents concept fusion, novel architectures and training paradigms could also be developed to selectively prevent unwanted fusion.

Existing work. Prior work includes retrieval-based architectures for revocable sensitive data access [46] and classifiers that self-identify concepts to generate each output [47], [48]. Recent work also explores concept-based explainability for text-to-image models via sparse autoencoders [49], [50].

Limitations. Most concept-based explainability work requires accessing and generating images to identify interpretable features. Mechanisms relying on isolated sensitive datastores face challenges in reliably classifying sensitive imagery (e.g. NSFW content), may make sacrifices in quality on benign concepts, and are further complicated by entangled concepts (e.g. children’s medical imagery).

3.2.1 Proof-of-Concept: Concept Fusion↩︎

A concern we raise in Open Problem 2 is that models may be able to combine high-risk concepts such as children and NSFW material. As shown in Figure 1, we find this to be the case with proxy concepts from CelebA [51]: we train diffusion models on separate images of people with blonde hair and eyeglasses, and find that they can generate the combined concept of blondes wearing eyeglasses with zero prior examples. Concurrent work from [45] similarly uses eyeglasses as a proxy for nudity, and finds that text-to-image models can generate children wearing eyeglasses even with 94% of child images removed. Developing alternative techniques to prevent concept fusion is thus an important area of future work. See App. 10.1.1 for full experiment details.

Open Problem A3: Resilience to harmful fine-tuning How can generative models be post-trained to prevent fine-tuning on CSAM or simultaneously fine-tuning on multiple CSAM-related concepts?

Unfortunately, even if base models appear ‘safe’, users can also unlock harmful capabilities through post-training procedures such as fine-tuning. Open-weight models allow users to adapt models on various tasks and are vulnerable to tampering [52]. CSAM perpetrators use GUI-based LoRA fine-tuning software like ComfyUI [53] or Ostris [54] to locally fine-tune open source models on CSAM [55]. “Nudifying” apps generating illegal deepfake sexual material of children similarly exploit the open source ecosystem, optimizing models for clothing removal and face-swapping [56]. Ideally, text-to-image models would allow benign fine-tuning while preventing these harmful uses.

Existing work. [57] propose self-destructing classifiers that degrade when fine-tuned on specific tasks. Similar approaches have been explored for diffusion models [58], [59].

Limitations. Solutions for self-destructing models that obstruct joint fine-tuning on separate concepts (e.g., adult sexual content and children) without blocking individual concept fine-tuning are largely unexplored. Similarly, strategies to prevent harmful LoRA fine-tuning (vs. full fine-tuning) are lacking. Methods that build resilience to harmful fine-tuning often require training on obstructed tasks; CSAM data access restrictions make this challenging.

We explore further problems in development, such as watermarking and minimizing human exposure, in App. 10.1.

a

b

Figure 1: Conditional diffusion models trained on images of blondes (top) and people in eyeglasses (middle) can generate blondes wearing eyeglasses (bottom), without overlapping training data. Fusion is effective with a critical threshold of training data (right)..

3.3 Call to Action↩︎

Below we outline take-aways for researchers, AI providers, and policymakers based on the identified open problems.

Research Directions

  • Assess partial data cleaning limits using proxy datasets; explore architectures preventing unwanted concept fusion for images/video.

  • Build self-destructing models targeting the “nudifying” ecosystem.

  • Create mechanisms to obstruct harmful fine-tuning that do not require access to the harmful material.

  • Develop robust content provenance: e.g. models that sharply degrade if finetuned to remove a watermark and localized provenance to handle inpainting.

Given data access restrictions, researchers and AI providers should also establish collaborations with reporting hotlines and LE to enable effective interventions and evaluation.

Concrete Steps for AI Providers

  • Engage CSAM survivors to understand their perspectives on partial data cleaning and re-victimization.

  • Partner with hotlines and LE for secure, scoped CSAM access to implement fine-tuning resilience.

  • Reliably label model outputs with C2PA or an equivalent standard and indelible watermarks. For open source models, prioritize tamper-proof content provenance solutions.

Policymakers play a critical role by establishing long-term pathways to encourage safer model designs. Regulatory efforts, research funding, and private-public partnerships that prioritize child safety are all important for effective action.

Policy Goals

  • Establish avenues for scoped, secure evaluation of AIG-CSAM capabilities by vetted institutions.

  • Resource standards organizations (e.g. NIST) to ensure that their guidance for content provenance in the open source setting remains up to date.

  • Instruct regulators to engage with the fine-tuning software ecosystem; assess what interventions are in place to prevent the misuse of these tools.

  • Task grant-making institutions to prioritize related research efforts in AI child safety.

4 Deployment Safeguards↩︎

Once AI models have been trained and evaluated for child safety, they should also be deployed with safeguards built into the release and distribution processes. Safeguards may include techniques such as content moderation, usage violation detection, reporting mechanisms, and transparency disclosures. These defenses are complicated by the AI model ecosystem, with providers building iteratively off each other’s systems via post-training and distillation.

4.1 Current Approaches↩︎

Closed-source AI providers currently use safety classifiers [60] to detect harmful inputs/outputs. They can direct users who attempt to generate illegal content to help resources [33]. Some open- and closed-source providers establish user reporting pathways for violating outputs, prompts, and models. Open-source models may also include safety classifiers [61], or be safety fine-tuned to reject adversarial prompts [62], [63].

To anticipate adversaries who may attempt to circumvent safeguards, providers employ red teaming to identify harmful prompts for training classifiers [35]. However, this is less common on third-party platforms [33]. While platforms like HuggingFace generally prohibit third-party models with AIG-CSAM capabilities, enforcement of these policies varies [64]. One measure for developer transparency and accountability is recommending model cards.

4.2 Open Problems↩︎

Adversaries can jailbreak systems by disguising harmful signals [65], [66]; text-to-image models are particularly vulnerable [67][70].

Open Problem B1: Effective prompt/output detection How can harmful prompts/outputs be reliably detected, in an adversarial ecosystem with limited access to in-distribution examples?

Open source deployments face additional challenges: content moderation is not feasible, safety filters are vulnerable [71], and fine-tuning interventions are easily reversed [72], [73].

Existing work. Existing harmful prompt refusal requires instruction fine-tuning on similar prompts or few-shot examples in context [69], [74]. Other solutions include automated red teaming tools to search for violatory prompts [75], zero-shot prompt classifiers [60] and guideline based model training to reject unsafe prompts [76].

Limitations. Red teaming prompts and guidelines may not reflect actual CSAM offender behavior. Most organizations do not have access to CSAM; accurate classifiers require offender data. For those that do, there is no public benchmark to evaluate their solutions. Research to strengthen solutions in the open source setting is broadly lacking.

Open Problem B2: Automated model assessment How can models be assessed for AIG-CSAM capabilities and CSAM training data automatically?

Third-party models constitute a significant chunk of the model ecosystem, and are not always assessed for CSAM risks pre-deployment. Scalable assessment of deployed models is needed to enforce platform policies and prevent distribution of models with AIG-CSAM capabilities.

Existing work. Training data extraction techniques [77] have been applied to detect the presence of specific media in training data. Data attribution techniques for generative models could also prove useful to identify the samples that enable AIG-CSAM [78][80]. The nascent area of mechanistic interpretability (MI) aims to examine learned weights to understand large models [81], with techniques such as automated circuit discovery attempting to identify computational subgraphs that implement specific model behaviors [82], allowing auditing of specific capabilities by searching relevant subgraphs.

Limitations. Research using MI to audit CSAM generation capabilities is unexplored; seeking out CSAM trained models to assess such techniques may open researchers to risk [83]. Current techniques for training data extraction rely on prompting for AIG-CSAM, which violates US law. Text-to-video assessment is also broadly unexplored.

We explore further problems in deployment, such as model transparency and standardize assessments, in App. 10.2.

4.3 Call to Action↩︎

As with the previous open problems, in the absence of collaborations that enable researchers access to CSAM or CSAM trained models, there are still lines of research that can be pursued that directly ladder up, e.g.:

Research Directions

  • Design image-free auditing to enable upstream detection (e.g. auditing before diffusion denoising completion [84], [85].)

  • Harden models to white-box attacks using natural-language safety specifications [76].

  • Develop training data extraction techniques that do not rely on direct prompting or model training

  • Explore concept fusion via MI to identify and downweight neurons enabling adult-child combinations.

AI Providers have unique visibility into how offenders are misusing their platforms. They should prioritize efforts to use and share this data and other related metadata available to strengthen safeguards, including efforts such as:

AI Providers

  • Join data sharing programs (e.g. [86]) with trusted organizations to broaden access to adversarial prompts across platforms.

  • Deploy metadata and other signals to detect policy-violating models. Resource teams to meet the volume of models uploaded/downloaded.

  • Partner with external teams pre-deployment for model evaluation, sharing platform-specific offender behavior data for effective assessment

Policymakers are key to establishing broader systems and structures for building public trust in safety technologies used to safeguard generative models. They also play a critical role in drawing clear lines in the sand that allow for scalable, preventative efforts.

Policymakers

  • Resource institutions like NIST to establish pathways for public benchmarking of safety tech.

  • Pass legislation creating liability for intentional development or distribution of models built to produce AIG-CSAM or “nudify” pictures of children.

  • Require that developers disclose whether they conducted CSAM filtering on their training data (e.g. [87], [88])

5 Safe Model Maintenance↩︎

Finally, model monitoring and maintenance is necessary to address emerging and evolving threats to children, maintain efficacy against an evolving technology stack, and reflect the broad nature of the ecosystem (where interweaving technology providers and systems may inadvertently reinforce harms, or create new harms [89]). In particular, open source models are highly vulnerable to abliteration: fine-tuning to remove safety guardrails [90], allowing offenders to enhance CSAM generation or “nudify” minors. Technology aggregators (model hosting platforms, app stores, search engines, AI character platforms, etc.) enable easy discovery of these harmful models [91], [92].

5.1 Current Approaches↩︎

Current industry safety solutions for detecting harmful models are primarily reactive—removing flagged models rather than evaluating before providing access. While “nudifying” applications can be easily discovered via simple text matching searches [93], interventions to delist or suppress such sites and resources are largely lacking. Those that do incorporate detection strategies report using model hashlists (blocking uploads of files that are on hashlists of known CSAM optimized models [33], [94]) or similar efforts such as removing advertisements of “nudifying” applications [95]. Industry further reports efforts to maintain the quality of their own detection technologies, include child safety policies for each of their services, and collaborate with child safety organizations [33].

5.2 Open Problems↩︎

Open Problem C1: Identifying abliterated models How can we identify models and services that have been optimized for CSAM and “nudification”?

With thousands of models, apps, and services created and uploaded daily, rapid identification of those optimized for harmful purposes remains a critical gap.

Existing work. Model fingerprinting uses adversarial attacks to compare outputs between original and suspected stolen models for IP protection [96]. Model diffing exploits mechanistic differences between base and fine-tuned models to identify model-specific concepts [97], [98]. AIG-CSAM model hashlists can be built using cryptographic hashing. Limitations. Model fingerprinting requires insight into the “original” model, which is challenging to determine for models optimized by malicious actors. Model diffing is not well explored for text-to-image models, particularly for LoRA fine-tuning (often used by CSAM perpetrators). Cryptographic hashing detects exact model replicas but not minor modifications. Further, sourcing such hashlists can require accessing Tor onion sites dedicated to child abuse, which comes with significant legal and wellness risks.

Open Problem C2: Robust unlearning How can we reliably erase the concept of CSAM from generative models?

CSAM might be found in training data after a model is deployed; concept fusion opens up other pathways for CSAM generation. The gold standard is to retrain the model from scratch without harmful data, but retraining can be prohibitively expensive and time-consuming [99]. If the model was built by a different developer than the person who discovered the issue, engaging the original developer to conduct a full re-train may be challenging.

Limitations. Existing work has extensively explored approximate machine unlearning or concept erasure for text-to-image models, to remove knowledge with only a textual description of the harmful concept [63], [100]. However, these methods are vulnerable to adversarial prompts [101] and only provide probabilistic guarantees that a model no longer retains certain knowledge. CSAM requires exact unlearning, providing strict guarantees that CSAM images have no effect on model output [99], [102].

Open Problem C3: Protecting user imagery How do we proactively protect users’ imagery from unwanted AI-generated manipulation?

Solutions at the model level to prevent adversarial optimization are necessary, but do not afford end users (or platforms hosting user generated content) agency to proactively protect their own content from unwanted AI-manipulation. Existing work. Research exploring image immunization [103] involves injecting imperceptible perturbations into the image, such that image editing software fails to successfully edit the image. Some recent efforts specifically focus on protecting children’s imagery [104]. In the IP protection space, similar solutions have been proposed to disrput a model’s ability to mimic the style of particular artists [105].

Limitations. Some research indicates that these image perturbation strategies are not robust to simple attacks such as image upscaling [106]. Further, solutions to protect video imagery are lacking. Effectively evaluating these techniques ability to protect children’s imagery from “nudification” requires attempting to generate such imagery, which has ethical and legal implications.

Open Problem C4: Securing AI agents How can we prevent the misuse of AI agents and code generation to facilitate child sexual abuse?

Although criminal actors do not yet appear to be adopting automated code generation and AI agents for child sexual exploitation, the tools are already being used in other criminal enterprises, such as malware creation [107] and building hidden webcam recording software [108]. AI agents capable of relationship building could enable sexual extortion schemes; code generation could automate the creation of “nudifying” software, even if prebuilt nudifying apps are banned.

Existing work. Safety for agentic systems is an emerging field [109]. Current work focuses on early detection and prevention of misuse. As with traditional red teaming and filtering, most solutions rely on training models using examples of prior misuse, such as AI-generated malware [110][112]. Such examples are legal to develop and train on, and leading labs actively build this data.

Limitations. Developing “nudification” code or sexual extortion prompts is more ethically ambiguous than malware, raising data legality and researcher wellness concerns similar to directly red teaming for CSAM. Robust evidence that these emerging technologies are being misused in child safety contexts may be a prerequisite for industry prioritization, but obtaining the visibility needed to build such evidence remains challenging.

We discuss additional open problems in model maintenance, such as safeguard assessments, in App. 10.3.

5.3 Call to Action↩︎

Interventions for safe model maintenance are unique in that they can be both cross-platform and model specific, requiring research that covers broad surface areas for harm.

Research Directions

  • Build solutions to detect “nudifying” applications.

  • Develop model hashing techniques that are robust to minor modifications of the model.

  • Explore strategies that provide strong guarantees when remediating/unlearning harmful models.

  • Establish image/video protection techniques for chidren’s imagery, using proxy data and concepts.

AI Providers are already positioned to engage in both model specific and cross-platform safeguarding efforts:

AI Providers

  • Partner with cross-industry child safety organizations (e.g. the Tech Coalition) to enable cross-platform monitoring and measurement.

  • Partner with user-generated content platforms to evaluate strategies for protecting user imagery from unwanted AI-generated manipulation originating from your foundation models

  • Model hosting platforms jointly establish cross-industry consistency in policies and enforcement for identifying & removing harmful third-party models

Policymakers can create incentives and pathways such that stakeholders invest in cross-platform safety:

Policymakers

  • Task regulators to review “nudifying” services (websites, apps, etc.) for unfair, deceptive and fraudulent business practices that promote AIG-CSAM creation and distribution.

  • Mandate child safety specific impact and risk assessments for AI systems before deployment.

  • Direct regulators to assess platform (e.g. model hosting companies, search engines and app stores) strategies for preventing the distribution of models, apps and services that enable CSEA.

6 Alternative Views↩︎

The most immediate alternative perspective is that child safety does not require fundamentally new AI safety approaches, but rather stronger or more consistently applied variants of existing approaches. Indeed, the core technical challenges (e.g., dataset governance, robust content filtering, adversarial testing, and post-deployment monitoring) are not unique to child safety and are broadly relevant to problems like disinformation, non-consensual imagery, and extremist content. Our work argues that despite the seeming similarities to these problems, the legal and ethical restrictions surrounding AIG-CSAM present non-trivial additional hurdles to existing approaches, necessitating in some cases entirely new techniques. While the problems identified are necessary to reduce the harms of AIG-CSAM, we believe many of them could prove beneficial for these other, less regulated domains.

Another view is that the problem of AIG-CSAM simply cannot be solved, or that we should instead focus AI safety efforts on other, existential risks [113]. Others may go further to suggest that AIG-CSAM is not harmful or is at least preferable to non-AI generated CSAM. Here, we note the significant evidence to the contrary [114], and advocate for a balanced position to invest efforts in multiple safety areas, especially as we could potentially make immediate headway on reducing the harms of AIG-CSAM during a time when reports are increasing sharply [2], [5], [6].

Finally, one may question whether technical AI safety interventions should be the primary locus of response at all. We could instead emphasize social, legal, and preventive measures such as digital literacy education for children and parents, victim support services, updated criminal law and international cooperation, and platform liability regimes that deter harmful deployment. From this standpoint, overemphasis on building new AI-specific safeguards may create a false sense of technological solvability while underinvesting in root causes. Others may go further to argue that policy solutions would eliminate the need for developing new technical approaches altogether, reducing the legal barriers to applying existing safety approaches. Our position reflects that policy solutions must come paired with, and be informed by, technical advancements.

7 Conclusion, Limitations & Future Work↩︎

In this work we argue that current safety techniques assume data accessibility, transparency, and evaluation practices that are in many cases incompatible with ethical and legal constraints surrounding CSAM. We then propose targeted recommendations and a set of 15 open technical problems that can improve child safety in the AI ecosystem.

We note limitations and areas for future work in our study. We focus on AI-generated imagery. However, exploitation spans multiple modalities, including text sexualizing minors [115], generative chatbots using celebrity voices for sexual conversations with children’s accounts [116], and video deepfakes of child sexual abuse [117]. While some of the issues identified in this work may also translate to other modalities, we leave the identification of technical open problems in non-image modalities as future work.

We also note that the risks and open problems identified throughout are primarily from a U.S.-centric perspective. The policy institutions, regulatory bodies, and frameworks referenced are also primarily U.S.-centric. Global regulation in this space is actively evolving [118], [119], and those evolutions may add additional nuance and context to the content in this work. A short discussion of policy tradeoffs is provided in Appendix 11, but exact policy implementations are left for future work and discussion with policymakers.

Moving forward, measuring and reducing the utility cost of existing safety methods is critical for adoption. At present, most AI-generated abusive material and models optimized for AIG-CSAM derive from open-source models. As such, the most pressing open problems are those that address the adversarial misuse of open-source models. Resilience to harmful fine-tuning (A3) is imperative, as perpetrators actively use GUI-based LoRA tools to abliterate models for AIG-CSAM generation. Other urgent problems include watermark robustness (A5), moderating the hobbyist ecosystem (B5), detecting abliterated models (C1), and robust unlearning (C2); we direct readers to [52] on broader open-source safety risks. Establishing collaborative venues (e.g., conference workshops) that foster communication between stakeholders could accelerate progress on these open problems.

Acknowledgments↩︎

We thank Steven Wu for his contributions in early planning for this work, and Stephen Casper for his feedback on later drafts. We also thank Tim O’Gorman, James Williams, Jim Pitkow, and Melissa Stroebel from Thorn for their reviews and comments on the paper.

8 Glossary↩︎

Table 2: Key definitions used throughout this paper.
CSEA Child sexual abuse and exploitation
AIG-CSAM AI-generated child sexual abuse material
Reporting hotlines National Center for Missing & Exploited Children (NCMEC), Internet Watch Foundation (IWF)
AI Developer The entity responsible for designing, training, and testing an AI model or system.
AI Deployer The entity that puts an AI system into service and controls its operation, often by integrating it into a product or platform.
AI Provider Any entity that makes an AI system or model available for use.
Table 3: Key challenges and broad restrictions across the proposed open problems in preventing CSAM generation.
Open Problem # 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15
Data restrictions
Evaluation restrictions
Adversarial environment
Wellness implications

9 Additional Related Work↩︎

Extending AI safety to AIG-CSAM prevention naturally builds on diverse prior work that mitigates AI risks. Throughout this work, we highlight many relevant papers in these domains, e.g., on topics such as on red teaming [35], model fingerprinting [96], and tamper resistance [57]. In addition, several papers (1) outline broader open problems in AI safety, or (2) propose child-safety specific mitigations. We highlight these areas of related work here.

9.0.0.1 Open Problems in AI Safety.

[52] outline open problems in derisking open-source models, several of which directly address child safety, such as reliable tamper resistance (Open Problem 3: Resilience to harmful fine-tuning). Similarly, [117] demonstrate problems with the open-source model landscape that enable video deepfake propagation, including video CSAM. Specific to Open Problem 1 (Partial data cleaning), [45] find that partial data filtering is insufficient for CSAM prevention and deconstruct open subproblems, such as accurate child detection.

9.0.0.2 Child Safety Mitigations.

Some existing works specifically address open problems outlined in this paper for preventing AIG-CSAM. For example, [30] propose annotation pipelines to improve data cleaning (Open Problem 1) and apply them to Stable Diffusion. Work from [120] develops an end-to-end CSAM classifier and proposes an evaluation methodology incorporating proxy data and controlled access to CSAM (Open Problem 6). There remains a significant need for further work on these problems.

9.0.0.3 Comparison to Other Restricted Domains.

Child sexual abuse material (CSAM) is unique in that it is the only media in the United States that is illegal to create, possess, or distribute [121]. This is also broadly true under global regulatory regimes [122]. While there are other domains with data that may be similarly traumatic to view, CSAM is unique in its illegal status. Government-sensitive material bears some similarity in that its distribution is limited [123], but CSAM access is not tied to security clearance levels the way other government-sensitive data is. Medical records under HIPAA also bear similarity in that data is restricted [124], but they are not subject to evaluation, generation, and wellness constraints that CSAM falls under.

10 Additional Open Problems↩︎

10.1 Developing Safe Models by Design↩︎

Open Problem A4: Minimizing human exposure How does exposure to AIG-CSAM affect red teamers? How can human exposure be minimized?

The emotional and mental toll of CSAM exposure is well documented across moderation [125], law enforcement work [126], and data labeling [127], with additional wellness implications for red teamers [128]. These harms demonstrate the clear need for red teaming solutions that minimize human exposure to CSAM.

Existing work. Several frameworks for automated red teaming text-to-image models have been proposed [75], [129], [130], and existing industry content review tools incorporate wellness features like image blurring [131]. Limitations. Text-to-image models are particularly vulnerable to out-of-distribution prompts, requiring robust human red teaming [67], [132]. Offenders quickly discover new jailbreaking mechanisms, making it challenging to ensure testing remains relevant. Moreover, evaluating wellness features for red-teamers requires sociotechnical studies with industry and red teaming service providers, who may not have incentive to participate in such studies.

Open Problem A5: Watermark robustness How can text-to-image watermarks be made robust to removal, spoofing, and partial edits?

In the CSAM space, offenders actively combat safety interventions, including stripping metadata and other identifying factors in the image (e.g. watermarks). Researchers have explored robust watermarking techniques that are less susceptible to spoofing or removal [133], [134]. Methods like Stable Signature incorporate the watermark directly in the model weights during pretraining [41].

Limitations. Many schemes remain vulnerable to simple attacks [135]. Watermarks can be removed by partially modifying an image [136] or editing code in image generation scripts [137]. Watermarking techniques that are more robust to fine-tuning and erasure tend to be more vulnerable to being stolen and spoofed on other images [135]. Robust solutions that include a history of partial edits are also lacking.

10.1.1 Experiment Details: Concept Fusion↩︎

To demonstrate concept fusion, we trained conditional flow matching models on 128x128 CelebA images. Each model comprised a 295M parameter UNet and was trained for 700 epochs. Given two attribute classes \(A, B\) (e.g., blonde hair, eyeglasses), models were trained on images from \(A \setminus B \cup B \setminus A\)—for instance, blondes without eyeglasses and non-blondes with eyeglasses (see Section 3.2.1 for examples). We train on 4K images of a certain hair color (blonde, black) without eyeglasses and a varying number of images (64-4K) without that hair color but with eyeglasses.

We measured the models’ propensity to generate compositional images from \(A \cap B\) (blonde hair and eyeglasses in Figure 1, black hair and eyeglasses in Figure 2) while varying dataset composition. In Figure 1, we test two generation methods of (1) unconditional generation or (2) conditioning on an average class vector. We then take the maximum detection rate over these two methods. We observe a sharp threshold in the number of samples required for composition: with 750 eyeglasses samples, composition likelihood was \({\sim}0.1\%\); with 1000 samples, this increased \(15\times\) to \(1.5\%\).

Varying the sampling strategy can further increase this ratio. In Figure 2, we evaluate a sequential conditioning approach: instead of denoising unconditionally for \(T\) steps, we first denoise for \(n < T\) steps conditioned on class \(A\), then for \(T - n\) steps conditioned on class \(B\). While a fixed conditioning (unconditional or average) measures propensity, this strategy measures capability. Here, capability proves substantially stronger—the model generates images from \(A \cap B\) over \(25\%\) of the time with just 250 eyeglasses samples (\(6.2\%\) of training data).

Figure 2: We test composition of the black hair and eyeglasses concepts, varying the number of eyeglasses samples from 64 to 4K. When increasing from 64 to 250 samples, capability to produce compositional images triples from 9% to 27%.

Our findings align with [31], who also observe a sharp sample threshold for concept composition. This suggests partial data cleaning may be effective, though the bar is high: in our setting, just 64 images of a “cleaned concept” still enable composition \({\sim}10\%\) of the time. Future work should explore how this threshold scales with larger diffusion models. Concurrent work from [45] more directly evaluates model capability to compose children with other concepts and highlights additional challenges such as child detection. We note both directions as ways to pursue child-safety research on the open problems posed here, at different levels of abstraction.

10.2 Deployment Safeguards↩︎

Open Problem B3: Standardized safety assessments How can we standardize assessments for AIG-CSAM capabilities?

Standardized safety assessments allow for consistent and transparent model evaluation. For child safety in particular, building confidence and trust in evaluations requires assurance that the assessment is robust, and not unduly influenced by other incentives, e.g., product deadlines.

Existing work. External red teaming and benchmarking are standard for assessing generative models. External domain expertise can help discover novel issues [138]. Benchmarking supports scalable reproducibility, allowing multiple models from different developers to undergo the same evaluation [139].

Limitations. Safety benchmarking assessments may correlate with model capabilities rather than actual safety [140]. Other studies highlight fundamental gaps in AI safety assessments, particularly for non-text modalities [141]. Given the adversarial nature of CSAM offenders and the wellness and legal barriers, static benchmarks quickly becomes outdated as offenders develop new strategies for generating AIG-CSAM.

Open Problem B4: Model transparency How can we use model cards to encourage transparency without inadvertently enabling offenders?

Model cards are the industry standard for disclosing information about models. For AIG-CSAM, documenting child safety interventions creates a natural pause point for developers to assess safeguards.

Existing work. Model cards are intended to provide fair assessment on a variety of human critical factors, e.g., bias [142]. Model card format can influence how interpretable the information is to non-technical audiences [143], enabling ethical decision-making for laypersons as well. Even the act of filling out model cards can elicit further ethical consideration from participating developers [144].

Limitations. Disclosing safety interventions is only effective if deployment is actually contingent on implementing them; currently, models without CSAM safeguards can still be released. While transparency enables good-faith actors to identify well-safeguarded models, it also enables offenders to discover vulnerable ones.

10.2.1 Proof-of-Concept: Model Cards↩︎

As a second proof-of-concept experiment, we manually audit public model and system cards for 14 widely used image and video generators (Table 4), asking whether they document child-safety relevant risks. These 14 models were chosen based on their presence in popular video generation [145] and image generation leader boards [146]. This analysis is limited to publicly available documentation and does not make assessments about any deployed safeguards, absence of documentation should not be interpreted as absence of safety measures. We mark whether a public card exists and whether it reports any evaluation targeted at children, explicit sexual content in general, and CSAM.

Table 4: Presence of child safety related disclosures in public documents for selected image and video generation models. This analysis examines disclosures regarding the categories of children, explicit content, and CSAM and is limited to publicly available documentation; the absence of documentation does not necessarily imply the absence of safeguards.
Model
Exists Child Explicit CSAM
Sora
Sora 2
DALLE-3
Imagen 4
Veo 3.1
FLUX.1
DeepSeek Janus
SDXL
Z-Image
RunwayML Gen 4.5
Grok Imagine
Midjourney
Pika
Luma

Among the systems with model or system cards, only OpenAI’s Sora models include dedicated sections that explicitly discuss child safety and CSAM. Several of the most widely deployed systems (e.g., Midjourney, Grok Imagine, Pika, Luma) have no public model card at all, despite their scale and popularity.

Open Problem B5: Hobbyist ecosystem How can we encourage adoption of safety best practices across the diverse ML hobbyist ecosystem?

AI safety efforts typically target industry providers rather than hobbyists who build on foundation models [55]. This leaves a gap: downstream actors may lack the resources, incentives, or oversight to maintain the CSAM safeguards built into the original model.

Existing work. AI developers broadly recognize ethical dilemmas but often lack resources and training to navigate them [147]. Education-based prevention of child sexual abuse is well-studied and established as a standard approach [148], [149]. Research on hobbyist developers highlights intellectual stimulation as a primary motivation [150].

Limitations. Education efforts to promote developer awareness and CSAM prevention practices remain understudied. Researching those online communities carries risks of harassment and doxing [151], compounded by the high-stakes nature of child sexual abuse—a topic that often prompts defensiveness and avoidance when developers confront unintended consequences of their models.

10.3 Safe Model Maintenance↩︎

Open Problem C5: Assessing safeguards How do third-party auditors and users effectively assess the efficacy of implemented safeguards?

Even where safeguards have been implemented or reported as implemented, assessing their effectiveness—individually and within the broader system context—is necessary for building trust and transparency.

Limitations. As noted earlier, most AI safety assessments focus on individual models through red teaming or benchmarks. Mechanisms for assessing complex AI systems are lacking. Sociotechnical assessments that account for how users actually engage with platforms, including offender behavior, platform-specific risks, and cross-platform dynamics, remain uncommon.

Companies may lack incentive to provide access necessary for these studies, which are important for building shared understanding and trust. Some metrics require cross-platform measurement, such as tracing circulating AIG-CSAM from downstream tools back to source models as a proxy for safeguard robustness.

11 Policy Tradeoffs↩︎

Enacting the policy recommendations highlighted in this work may require difficult decisions and challenging tradeoffs. For example: establishing secure pathways for vetted institutions to evaluate AIG-CSAM capabilities could improve accountability and benchmarking, but would also require carefully scoped governance structures to prevent misuse, leakage, or other harms associated to exposure to highly sensitive material. When considering regulatory tools for transparency: mandatory disclosures on CSAM filtering, child-safety risk assessments, and platform prevention strategies all require clear standards for compliance auditing. In the absence of these, tradeoffs between flexibility and enforceability may result in company disclosures that are overly vague and lack meaningful detail. Even with such standards in place, challenges emerge around establishing the right level of specificity within these standards. Narrow requirements risk becoming quickly outdated as AI systems evolve, while overly broad requirements could create inconsistent enforcement.

These examples do not reflect the full scope of tradeoffs and decisions that may need to occur when enacting policy solutions in this space. Even so: we emphasize that, while these decisions are important and challenging, not enacting policy solutions is just as much a decision. The outcome of that decision, when it comes to AIG-CSAM, is tangibly apparent: more victims and more harm. This outcome is unacceptable.

References↩︎

[1]
IWF, “How AI is being abused to create child sexual abuse imagery.” https://www.iwf.org.uk/about-us/why-we-exist/our-research/how-ai-is-being-abused-to-create-child-sexual-abuse-imagery/; Internet Watch Foundation, 2023.
[2]
IWF, “What has changed in the AI CSAM landscape?” https://www.iwf.org.uk/media/nadlcb1z/iwf-ai-csam-report_update-public-jul24v13.pdf; Internet Watch Foundation, 2024.
[3]
G. Paltieli and G. Freud, “How predators are abusing generative AI.” ActiveFence, 2023, [Online]. Available: https://www. activefence.com/blog/predators-abusing-generative-ai.
[4]
D. Thiel, M. Stroebel, and R. Portnoff, “Generative ML and CSAM: Implications and mitigations,” Stanford Digital Repository, 2023.
[5]
P. Davis, “The deepfake dilemma: New challenges protecting students, confidentiality.” https://www.missingkids.org/blog/2025/the-deepfake-dilemma-new-challenges-protecting-students-confidentiality; National Center for Missing & Exploited Children (NCMEC), 2025.
[6]
IWF, “Full feature-length AI films of child sexual abuse will be ‘inevitable’ as synthetic videos make ‘huge leaps’ in sophistication in a year.” https://www.iwf.org.uk/news-media/news/full-feature-length-ai-films-of-child-sexual-abuse-will-be-inevitable-as-synthetic-videos-make-huge-leaps-in-sophistication-in-a-year/; Internet Watch Foundation, 2025.
[7]
Thorn, “Sexual extortion & young people: Navigating threats in digital environments.” https://info.thorn.org/hubfs/Research/Thorn_SexualExtortionandYoungPeople_June2025.pdf, 2025.
[8]
H. Wolbers et al., “Sexual extortion of australian adolescents: Results from a national survey,” Trends and Issues in Crime and Criminal Justice, 2025.
[9]
Thorn, “Youth perspectives on online safety.” https://info.thorn.org/hubfs/Research/Thorn_23_YouthMonitoring_Report.pdf; Thorn, 2024.
[10]
S. Kapoor et al., “On the societal impact of open foundation models,” in International conference on machine learning, position paper track, 2024.
[11]
Thorn and All Tech is Human, “Safety by design for generative AI: Preventing child sexual abuse.” https://info.thorn.org/hubfs/thorn-safety-by-design-for-generative-AI.pdf, 2024.
[12]
A. Birhane, V. Prabhu, S. Han, V. Boddeti, and S. Luccioni, “Into the laion’s den: Investigating hate in multimodal datasets,” Advances in Neural Information Processing Systems, Track on Datasets and Benchmarks, 2023.
[13]
M. Musser, “A cost analysis of generative language models and influence operations,” arXiv preprint arXiv:2308.03740, 2023.
[14]
A. Peppin et al., “The reality of ai and biorisk,” in ACM conference on fairness, accountability, and transparency, 2025.
[15]
S. Hazra, B. P. Majumder, and T. Chakrabarty, “Position: AI safety should prioritize the future of work,” in International conference on machine learning, position paper track, 2025.
[16]
J. Hazell, “Spear phishing with large language models,” arXiv preprint arXiv:2305.06972, 2023.
[17]
FBI, “Criminals use generative artificial intelligence to facilitate financial fraud.” https://www.ic3.gov/PSA/2024/PSA241203; Federal Bureau of Investigations (FBI), 2024.
[18]
Thorn, “Thorn and all tech is human forge generative AI principles with AI leaders to enact strong child safety commitments.” https://www.thorn.org/blog/generative-ai-principles/; Thorn, 2024.
[19]
J. Vanian, “Meta files lawsuit against developer of CrushAI ‘nudify’ app,” CNBC, 2025, [Online]. Available: https://www.cnbc.com/2025/06/12/meta-files-lawsuit-against-developer-of-crushai-nudify-app.html.
[20]
US Congress, “TAKE IT DOWN act.” S.146, 2025, [Online]. Available: https://www.congress.gov/bill/119th-congress/senate-bill/146.
[21]
Federal Trade Commission, Accessed: 2025-11-26“Children’s online privacy protection rule ("COPPA").” Federal Trade Commission, 2025, [Online]. Available: https://www.ftc.gov/legal-library/browse/rules/childrens-online-privacy-protection-rule-coppa.
[22]
GDPR.eu, “What is GDPR, the EU’s new data protection law?” Proton AG, 2018, Accessed: Nov. 26, 2025. [Online]. Available: https://gdpr.eu/what-is-gdpr/.
[23]
US Congress, “Certain activities relating to material involving the sexual exploitation of minors.” 18 U.S.C. §2252, 2011, [Online]. Available: https://uscode.house.gov/view.xhtml?req=granuleid:USC-prelim-title18-section2252&num=0&edition=prelim.
[24]
US Congress, “Prosecutorial remedies and other tools to end the exploitation of children today act.” S. 151, 2003, [Online]. Available: https://www.congress.gov/bill/108th-congress/senate-bill/151.
[25]
R. Spence, A. Harrison, P. Bradbury, P. Bleakley, E. Martellozzo, and and Jeffrey DeMarco, “Content moderators’ strategies for coping with the stress of moderating content online,” Journal of Online Trust and Safety, vol. 1, p. 5, 2023.
[26]
U.S. Senate, REPORT Act.” S.474 - 118th Congress (2023-2024), 2024, [Online]. Available: https://www.congress.gov/bill/118th-congress/senate-bill/474.
[27]
S. Jasper, “How we detect, remove and report child sexual abuse material.” https://blog.google/technology/safety-security/how-we-detect-remove-and-report-child-sexual-abuse-material/; Google, 2022.
[28]
R. Bommasani et al., “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, 2021.
[29]
S. S. A. Magid et al., “Is what you ask for what you get? Investigating concept associations in text-to-image models,” Transactions on Machine Learning Research, 2025.
[30]
D. Thiel, “Identifying and eliminating csam in generative ml training data and models,” Stanford Internet Observatory, Cyber Policy Center, 2023.
[31]
M. Okawa, E. S. Lubana, R. Dick, and H. Tanaka, “Compositional abilities emerge multiplicatively: Exploring diffusion models on a synthetic task,” Advances in Neural Information Processing Systems, 2023.
[32]
D. Zhang, K. Ahuja, Y. Xu, Y. Wang, and A. Courville, “Can subnetwork structure be the key to out-of-distribution generalization?” in International conference on machine learning, 2021.
[33]
R. Portnoff and M. Simpson, Design & Publication by Yena Lee, Cassie Coccaro, and Justus Hyatt“Safety by design: Annual progress report (report #4: April 2024 to april 2025),” Thorn, 2025. [Online]. Available: https://info.thorn.org/hubfs/Thorn_SafetyByDesign_AnnualProgressReport_April2024-April2025.pdf.
[34]
H.-E. Lee, T. Ermakova, V. Ververis, and B. Fabian, “Detecting child sexual abuse material: A comprehensive survey,” Forensic Science International: Digital Investigation, vol. 34, p. 301022, 2020.
[35]
D. Ganguli et al., “Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned,” arXiv preprint arXiv:2209.07858, 2022.
[36]
Google, Accessed: 2023-10-27“Google’s AI red team: The ethical hackers making AI safer.” 2023, [Online]. Available: https://blog.google/technology/safety-security/googles-ai-red-team-the-ethical-hackers-making-ai-safer.
[37]
S. Grossman, R. Pfefferkorn, and S. Liu, Accessed: 2025-09-11“AI-generated child sexual abuse material: Insights from educators, platforms, law enforcement, legislators, and victims. Version 1.” Stanford Digital Repository, 2025, [Online]. Available: https://purl.stanford.edu/mn692xc5736/version/1.
[38]
Coalition for Content Provenance and Authenticity, Coalition for Content Provenance and Authenticity“C2PA: Advancing digital content transparency and authenticity.” https://c2pa.org/, 2024, Accessed: Nov. 26, 2025. [Online].
[39]
N. Yu, V. Skripniuk, S. Abdelnabi, and M. Fritz, “Artificial fingerprinting for generative models: Rooting deepfake attribution in training data,” in IEEE/CVF international conference on computer vision, 2021.
[40]
Y. Wen, J. Kirchenbauer, J. Geiping, and T. Goldstein, “Tree-ring watermarks: Fingerprints for diffusion images that are invisible and robust,” Advances in Neural Information Processing Systems, 2023.
[41]
P. Fernandez, G. Couairon, H. Jégou, M. Douze, and T. Furon, “The stable signature: Rooting watermarks in latent diffusion models,” in IEEE/CVF international conference on computer vision, 2023.
[42]
K. O’Brien et al., “Deep ignorance: Filtering pretraining data builds tamper-resistant safeguards into open-weight LLMs,” arXiv preprint arXiv:2508.06601, 2025.
[43]
P. Maini et al., “Safety pretraining: Toward the next generation of safe ai,” in Advances in neural information processing systems, 2025.
[44]
A. Nichol et al., “Glide: Towards photorealistic image generation and editing with text-guided diffusion models,” in International conference on machine learning, 2022.
[45]
A.-M. Cretu et al., “Evaluating concept filtering defenses against child sexual abuse material generation by text-to-image models,” arXiv preprint arXiv:2512.05707, 2025.
[46]
S. Min et al., “Silo language models: Isolating legal risk in a nonparametric datastore.” 2024.
[47]
X. Wang, H. Chen, S. Tang, Z. Wu, and W. Zhu, “Disentangled representation learning,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.
[48]
M. E. Zarlenga et al., “Concept embedding models: Beyond the accuracy-explainability trade-off,” Advances in Neural Information Processing Systems, 2022.
[49]
B. Tinaz, Z. Fabian, and M. Soltanolkotabi, “Emergence and evolution of interpretable concepts in diffusion models,” Advances in Neural Information Processing Systems, 2026.
[50]
V. Surkov et al., “One-step is enough: Sparse autoencoders for text-to-image diffusion models,” Advances in Neural Information Processing Systems, 2025.
[51]
Z. Liu, P. Luo, X. Wang, and X. Tang, “Deep learning face attributes in the wild,” in IEEE/CVF international conference on computer vision, 2015.
[52]
S. Casper et al., “Open technical problems in open-weight ai model risk management,” Social Science Research Network, 2025.
[53]
Comfy, Accessed: 2025-09-08ComfyUI | Generate video, images, 3D, audio with AI.” https://www.comfy.org/, 2025, Accessed: Sep. 08, 2025. [Online]. Available: https://www.comfy.org/.
[54]
Ostris, AI Toolkit: The ultimate training toolkit for finetuning diffusion models.” 2024, [Online]. Available: https://github.com/ostris/ai-toolkit.
[55]
Thorn, “Mitigating the risk of generative AI models creating child sexual abuse materials,” Partnership on AI, Case Study, Nov. 2024. [Online]. Available: https://partnershiponai.org/wp-content/uploads/2024/11/case-study-thorn.pdf.
[56]
M. L. Ding and H. Suresh, “The malicious technical ecosystem: Exposing limitations in technical governance of AI-generated non-consensual intimate images of adults,” arXiv preprint arXiv:2504.17663, 2025.
[57]
P. Henderson, E. Mitchell, C. Manning, D. Jurafsky, and C. Finn, “Self-destructing models: Increasing the costs of harmful dual uses of foundation models,” in AAAI/ACM conference on AI, ethics, and society, 2023.
[58]
H. Gao, T. Pang, C. Du, T. Hu, Z. Deng, and M. Lin, “Meta-unlearning on diffusion models: Preventing relearning unlearned concepts,” in IEEE/CVF international conference on computer vision, 2025.
[59]
J. Pan et al., “Leveraging catastrophic forgetting to develop safe diffusion models against malicious finetuning,” Advances in Neural Information Processing Systems, 2024.
[60]
H. Inan et al., “Llama guard: Llm-based input-output safeguard for human-ai conversations,” arXiv preprint arXiv:2312.06674, 2023.
[61]
CompVis, Accessed: 2025-03-21“Stable diffusion safety checker.” https://huggingface.co/CompVis/stable-diffusion-safety-checker, 2022.
[62]
C. Kim, K. Min, and Y. Yang, “Race: Robust adversarial concept erasure for secure text-to-image diffusion model,” in European conference on computer vision, 2024.
[63]
R. Gandikota, J. Materzynska, J. Fiotto-Kaufman, and D. Bau, “Erasing concepts from diffusion models,” in IEEE/CVF international conference on computer vision, 2023.
[64]
D. E. Harris and D. Willner, “Was an AI image generator taken down for making child porn?” IEEE Spectrum, 2024, Accessed: Nov. 26, 2025. [Online]. Available: https://spectrum.ieee.org/stable-diffusion.
[65]
J. Ma et al., “Jailbreaking prompt attack: A controllable adversarial attack against diffusion models,” in Findings of the association for computational linguistics, 2025.
[66]
A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” arXiv preprint arXiv:2307.15043, 2023.
[67]
J. Rando, D. Paleka, D. Lindner, L. Heim, and F. Tramèr, “Red-teaming the stable diffusion safety filter,” in NeurIPS 2022 ML safety workshop, 2022.
[68]
A. Robey, E. Wong, H. Hassani, and G. J. Pappas, “Smoothllm: Defending large language models against jailbreaking attacks,” in Transactions on machine learning research, 2025.
[69]
Z. Wei, Y. Wang, A. Li, Y. Mo, and Y. Wang, “Jailbreak and guard aligned language models with only few in-context demonstrations,” arXiv preprint arXiv:2310.06387, 2023.
[70]
N. Jain et al., “Baseline defenses for adversarial attacks against aligned language models,” arXiv preprint arXiv:2309.00614, 2023.
[71]
H. Zhuang, Y. Zhang, and S. Liu, “A pilot study of query-free adversarial attack against stable diffusion,” in IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 2385–2392.
[72]
X. Qi et al., “Fine-tuning aligned language models compromises safety, even when users do not intend to!” in International conference on learning representations, 2024.
[73]
S. Hu, Y. Fu, S. Wu, and V. Smith, “Unlearning or obfuscating? Jogging the memory of unlearned LLMs via benign relearning,” in International conference on learning representations, 2025.
[74]
T. Markov et al., “A holistic approach to undesired content detection in the real world,” in AAAI conference on artificial intelligence, 2023.
[75]
G. Li, K. Chen, S. Zhang, J. Zhang, and T. Zhang, “ART: Automatic red-teaming for text-to-image models to protect benign users,” Advances in Neural Information Processing Systems, 2024.
[76]
M. Y. Guan et al., “Deliberative alignment: Reasoning enables safer language models,” arXiv preprint arXiv:2412.16339, 2024.
[77]
N. Carlini et al., “Extracting training data from diffusion models,” in USENIX security symposium, 2023.
[78]
K. Georgiev, J. Vendrow, H. Salman, S. M. Park, and A. Madry, “The journey, not the destination: How data guides diffusion models,” arXiv preprint arXiv:2312.06205, 2023.
[79]
J. Lin, L. Tao, M. Dong, and C. Xu, “Diffusion attribution score: Evaluating training data influence in diffusion models,” in International conference on learning representations, 2025, vol. 2025, pp. 22962–22989.
[80]
X. Zheng, T. Pang, C. Du, J. Jiang, and M. Lin, “Intriguing properties of data attribution on diffusion models,” in International conference on learning representations, 2024, vol. 2024, pp. 18417–18452.
[81]
A. Templeton et al., “Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet,” Transformer Circuits Thread, 2024, [Online]. Available: https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html.
[82]
A. Conmy, A. Mavor-Parker, A. Lynch, S. Heimersheim, and A. Garriga-Alonso, “Towards automated circuit discovery for mechanistic interpretability,” Advances in Neural Information Processing Systems, 2023.
[83]
P. Thaker, N. Kale, Z. S. Wu, and V. Smith, “Membership inference attacks for unseen classes,” arXiv preprint arXiv:2506.06488, 2025.
[84]
X. Yuan, X. Ma, L. Guo, and L. Zhang, “What lurks within? Concept auditing for shared diffusion models at scale,” ACM Conference on Computer and Communications Security, 2025.
[85]
F. Li, M. Zhang, Y. Sun, and M. Yang, “Detect-and-guide: Self-regulation of diffusion models for safe text-to-image generation via guideline token optimization,” in IEEE/CVF conference on computer vision and pattern recognition, 2025.
[86]
Tech Coalition, Accessed: 2025-11-26“Lantern: Advancing child safety through signal sharing.” Technology Coalition, 2025, [Online]. Available: https://technologycoalition.org/programs/lantern/.
[87]
European Commission, Shaping Europe’s digital future“The general-purpose AI code of practice.” https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai, 2025, Accessed: Nov. 26, 2025. [Online].
[88]
[89]
Thorn and WeProtect Global Alliance, “Evolving technologies horizon scan: A review of technologies carrying notable risk and opportunity in the fight against technology-facilitated child sexual exploitation,” Thorn; WeProtect Global Alliance, Technical Report, Dec. 2024. [Online]. Available: https://info.thorn.org/hubfs/Research/Thorn_x_WPGA_EvolvingTechnologies_Dec2024.pdf.
[90]
P. Henderson et al., “Safety risks from customizing foundation models via fine-tuning,” Policy Brief. Stanford Human-Centered Artificial Intelligence, 2024.
[91]
C. Stokel-Walker, “Thousands of pedophiles are using jail-broken AI character chatbots to roleplay sexually assaulting minors,” Fast Company, 2025, [Online]. Available: https://www.fastcompany.com/91290478/graphika-report-ai-chatbots-role-playing-sex-with-minors.
[92]
Thorn, “Deepfake nudes & young people: Navigating a new frontier in technology-facilitated nonconsensual sexual abuse and exploitation.” Thorn; https://www.thorn.org/research/library/deepfake-nudes-and-young-people/, 2024, Accessed: Nov. 26, 2025. [Online].
[93]
C. Gibson et al., “Analyzing the AI nudification application ecosystem,” in USENIX security symposium, 2025.
[94]
W. Hawkins, B. Mittelstadt, and C. Russell, “Deepfakes on demand: The rise of accessible non-consensual deepfake image generators,” in ACM conference on fairness, accountability, and transparency, 2025.
[95]
Meta, “Taking action against nudify apps.” Meta Platforms, Inc., Jun. 2025, Accessed: Jun. 10, 2025. [Online]. Available: https://about.fb.com/news/2025/06/taking-action-against-nudify-apps/.
[96]
J. Guan, J. Liang, and R. He, “Are you stealing my model? Sample correlation for fingerprinting deep neural networks,” Advances in Neural Information Processing Systems, 2022.
[97]
Anthropic Interpretability Team, “Stage-wise model diffing.” 2024, Accessed: Oct. 02, 2025. [Online]. Available: https://transformer-circuits.pub/2024/model-diffing/index.html.
[98]
J. Minder, C. Dumas, C. Juang, B. Chughtai, and N. Nanda, “Overcoming sparsity artifacts in crosscoders to interpret chat-tuning,” Advances in Neural Information Processing Systems, vol. 38, pp. 106423–106474, 2026.
[99]
A. F. Cooper et al., “Machine unlearning doesn’t do what you think: Lessons for generative AI policy, research, and practice,” in Advances in neural information processing systems, position paper track, 2025.
[100]
S. Lu, Z. Wang, L. Li, Y. Liu, and A. W.-K. Kong, “Mace: Mass concept erasure in diffusion models,” in IEEE/CVF conference on computer vision and pattern recognition, 2024.
[101]
Y. Zhang et al., “Defensive unlearning with adversarial training for robust concept erasure in diffusion models,” Advances in Neural Information Processing Systems, 2024.
[102]
L. Bourtoule et al., “Machine unlearning,” in IEEE symposium on security and privacy, 2021.
[103]
H. Salman, A. Khaddaj, G. Leclerc, A. Ilyas, and A. Madry, “Raising the cost of malicious ai-powered image editing,” International Conference on Machine Learning, 2023.
[104]
Australian Federal Police and Monash University, Accessed November 2025AFP and Monash University poison data to combat AI-generative crime.” Joint media release, 2025, [Online]. Available: https://www.afp.gov.au/news-centre/media-release/afp-and-monash-university-poison-data-combat-ai-generative-crime.
[105]
S. Shan, J. Cryan, E. Wenger, H. Zheng, R. Hanocka, and B. Y. Zhao, “Glaze: Protecting artists from style mimicry by text-to-image models,” in USENIX security symposium, 2023.
[106]
R. Hönig, J. Rando, N. Carlini, and F. Tramèr, “Adversarial perturbations cannot reliably protect artists from generative ai,” in International conference on learning representations, 2025.
[107]
Anthropic Threat Intelligence Team, “Threat intelligence report: August 2025,” Anthropic, Aug. 2025. Accessed: Nov. 26, 2025. [Online]. Available: https://www-cdn.anthropic.com/b2a76c6f6992465c09a6f2fce282f6c0cea8c200.pdf.
[108]
Google Threat Intelligence Group, “Adversarial misuse of generative AI.” Google Cloud Blog, 2025, [Online]. Available: https://cloud.google.com/blog/topics/threat-intelligence/adversarial-misuse-generative-ai.
[109]
D. Song, “Towards building safe & trustworthy AI agents and a path for science- and evidence-based AI policy.” Lecture slides for CS294/CS194-196: Large Language Model Agents, UC Berkeley, 2024, [Online]. Available: https://rdi.berkeley.edu/llm-agents/assets/dawn-agent-safety.pdf.
[110]
R. A. Popa and F. Flynn, “Introducing CodeMender: An AI agent for code security.” Google DeepMind Blog, 2025, [Online]. Available: https://deepmind.google/discover/blog/introducing-codemender-an-ai-agent-for-code-security/.
[111]
OpenAI, “Introducing aardvark: OpenAI’s agentic security researcher,” OpenAI, 2025, Accessed: Nov. 26, 2025. [Online]. Available: https://openai.com/index/introducing-aardvark/.
[112]
Anthropic, Accessed November 2025“Automate security reviews with Claude Code.” Claude Blog, Aug. 06, 2025, [Online]. Available: https://www.claude.com/blog/automate-security-reviews-with-claude-code.
[113]
D. Hendrycks, M. Mazeika, and T. Woodside, “An overview of catastrophic AI risks,” arXiv preprint arXiv:2306.12001, 2023.
[114]
C. Ó. Ciardha, J. Buckley, and R. S. Portnoff, “AI generated child sexual abuse material–what’s the harm?” arXiv preprint arXiv:2510.02978, 2025.
[115]
E. M. Cristina López G. Daniel Siegel, “Character flaws. School shooters, anorexia coaches, and sexualized minors: A look at harmful character chatbots and the communities that build them.” https://public-assets.graphika.com/reports/graphika-report-character-flaws.pdf; Graphika, 2025.
[116]
J. Horwitz, “Meta’s ‘digital companions’ will talk sex with users—even children,” Wall Street Journal. https://www.wsj.com/tech/ai/meta-ai-chatbots-sex-a25311bf, 2025.
[117]
M. Kamachee et al., “Video deepfake abuse: How company choices predictably shape misuse patterns,” arXiv preprint arXiv:2512.11815, 2025.
[118]
M. Clifton and T. Bristow, Accessed November 2025UK set to ban deepfake nudification apps,” Politico Europe, 2025, [Online]. Available: https://www.politico.eu/article/uk-set-to-ban-deepfake-nudification-apps-vawg/.
[119]
L. McMahon, “Deepfake ’nudify’ site fined £55,000 over lack of age checks,” BBC News, Nov. 2025, [Online]. Available: https://www.bbc.com/news/articles/cn8xq677l9xo.
[120]
W. Gutfeter, J. Gajewska, and A. Pacut, “Detecting sexually explicit content in the context of the child sexual abuse materials (CSAM): End-to-end classifiers and region-based networks,” in Joint european conference on machine learning and knowledge discovery in databases, 2023.
[121]
U.S. Department of Justice, Accessed: 2026-05-16“Citizen’s guide to U.S. Federal law on child pornography.” https://www.justice.gov/criminal/criminal-ceos/citizens-guide-us-federal-law-child-pornography, 2023.
[122]
International Centre for Missing & Exploited Children, Accessed: 2026-05-16“Child sexual abuse material: Model legislation & global review,” International Centre for Missing & Exploited Children, Alexandria, VA, Oct. 2023. [Online]. Available: https://cdn.icmec.org/wp-content/uploads/2023/10/CSAM-Model-Legislation_10th-Ed-Oct-2023.pdf.
[123]
W. J. Clinton, Accessed: 2026-05-16“Executive order 12968: Access to classified information.” 60 Fed. Reg. 40245, Aug. 02, 1995, [Online]. Available: https://www.archives.gov/isoo/policy-documents/eo-12968.html.
[124]
U.S. Congress, Accessed: 2026-05-16“Health insurance portability and accountability act of 1996.” Public Law 104–191, 110 Stat. 1936, Aug. 1996, [Online]. Available: https://www.govinfo.gov/app/details/PLAW-104publ191.
[125]
R. Spence, A. Bifulco, P. Bradbury, E. Martellozzo, and J. DeMarco, “The psychological impacts of content moderation on content moderators: A qualitative study,” Cyberpsychology: Journal of Psychosocial Research on Cyberspace, 2023.
[126]
J. R. Rimer, S. Brown, J. Martin, and A. Slane, ‘Once you see it you can’t unsee it’: Law enforcement trauma and immersion in child sexual abuse material,” Child Protection and Practice, 2025.
[127]
N. Rowe, “Kenyan moderators decry toll of training of AI models,” The Guardian, 2023, [Online]. Available: https://www.theguardian.com/technology/2023/aug/02/ai-chatbot-training-human-toll-content-moderator-meta-openai.
[128]
A. Q. Zhang et al., “The human factor in ai red teaming: Perspectives from social and collaborative computing,” in Conference on computer-supported cooperative work and social computing, 2024.
[129]
Z.-Y. Chin, C.-M. Jiang, C.-C. Huang, P.-Y. Chen, and W.-C. Chiu, “Prompting4debugging: Red-teaming text-to-image diffusion models by finding problematic prompts,” in International conference on machine learning, 2024.
[130]
B. Li et al., “DREAM: Scalable red teaming for text-to-image generative systems via distribution modeling,” arXiv preprint arXiv:2507.16329, 2025.
[131]
A. Das, B. Dang, and M. Lease, “Fast, accurate, and healthier: Interactive blurring helps moderators reduce exposure to harmful content,” in AAAI conference on human computation and crowdsourcing, 2020.
[132]
G. Daras and A. G. Dimakis, “Discovering the hidden vocabulary of dalle-2,” arXiv preprint arXiv:2206.00169, 2022.
[133]
W. Wan, J. Wang, Y. Zhang, J. Li, H. Yu, and J. Sun, “A comprehensive survey on robust image watermarking,” Neurocomputing, 2022.
[134]
S. Gowal et al., “SynthID-image: Image watermarking at internet scale,” arXiv preprint arXiv:2510.09263, 2025.
[135]
Q. Pang, S. Hu, W. Zheng, and V. Smith, “No free lunch in llm watermarking: Trade-offs in watermarking design choices,” Advances in Neural Information Processing Systems, 2024.
[136]
K. Tallam, J. K. Cava, C. Geniesse, N. B. Erichson, and M. W. Mahoney, “Removing watermarks with partial regeneration using semantic information,” arXiv preprint arXiv:2505.08234, 2025.
[137]
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in IEEE/CVF conference on computer vision and pattern recognition, 2022.
[138]
L. Ahmad, S. Agarwal, M. Lampe, and P. Mishkin, “OpenAI’s approach to external red teaming for AI models and systems,” arXiv preprint arXiv:2503.16431, 2025.
[139]
B. Vidgen et al., “Introducing v0. 5 of the ai safety benchmark from mlcommons,” arXiv preprint arXiv:2404.12241, 2024.
[140]
R. Ren et al., “Safetywashing: Do AI safety benchmarks actually measure safety progress?” Advances in Neural Information Processing Systems, 2024.
[141]
M. Rauh et al., “Gaps in the safety evaluation of generative AI,” Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 2024.
[142]
M. Mitchell et al., “Model cards for model reporting,” in Proceedings of the conference on fairness, accountability, and transparency, 2019.
[143]
A. Crisan, M. Drouhard, J. Vig, and N. Rajani, “Interactive model cards: A human-centered approach to model documentation,” in ACM conference on fairness, accountability, and transparency, 2022.
[144]
J. L. Nunes, G. D. Barbosa, C. S. De Souza, H. Lopes, and S. D. Barbosa, “Using model cards for ethical reflection: A qualitative exploration,” in Proceedings of the 21st brazilian symposium on human factors in computing systems, 2022.
[145]
Artificial Analysis, Accessed January 2026“Video generation arena leaderboard.” HuggingFace Spaces, 2025, [Online]. Available: https://huggingface.co/spaces/ArtificialAnalysis/Video-Generation-Arena-Leaderboard.
[146]
Artificial Analysis, Accessed January 2026“Text-to-image leaderboard.” HuggingFace Spaces, 2025, [Online]. Available: https://huggingface.co/spaces/ArtificialAnalysis/Text-to-Image-Leaderboard.
[147]
T. A. Griffin, B. P. Green, and J. V. Welie, “The ethical wisdom of AI developers,” AI and Ethics, vol. 5, no. 2, pp. 1087–1097, 2025.
[148]
S. K. Wurtele and M. C. Kenny, “Partnering with parents to prevent childhood sexual abuse,” Child Abuse Review: Journal of the British Association for the Study and Prevention of Child Abuse and Neglect, 2010.
[149]
A. Patterson, L. Ryckman, and C. Guerra, “A systematic review of the education and awareness interventions to prevent online child sexual abuse,” Journal of Child & Adolescent Trauma, 2022.
[150]
S. Koch and M. Kerschbaum, “Joining a smartphone ecosystem: Application developers’ motivations and decision criteria,” Information and Software Technology, vol. 56, no. 11, pp. 1423–1435, 2014.
[151]
P. Doerfler, A. Forte, E. De Cristofaro, G. Stringhini, J. Blackburn, and D. McCoy, I’m a professor, which isn’t usually a dangerous job’’: Internet-facilitated harassment and its impact on researchers,” Proceedings of the ACM on Human-Computer Interaction, 2021.