June 09, 2026
Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-generated child sexual abuse material, facilitate child sexual exploitation, and reduce barriers to harm. In this paper, we argue that protecting children from AI-facilitated sexual abuse requires new approaches to AI safety. Existing safety techniques assume data accessibility, transparency, and evaluation practices that are incompatible with the ethical and legal constraints surrounding child sexual abuse material. We examine how these constraints create new technical challenges, such as limitations on dataset auditing, red teaming, and fine-tuning prevention. In turn, we outline 15 open problems in online child sexual exploitation and abuse across the AI development lifecycle, from dataset curation and model design to deployment and long-term maintenance. We propose targeted recommendations for researchers, developers, and policymakers to bridge the gap between theoretical AI safety and the realities of child protection. Our work aims to reframe preventing AI-facilitated child sexual abuse as a central, safety-critical dimension for AI research, motivating work that translates responsible AI principles into concrete safeguards against the exploitation of children.
Artificial intelligence (AI) systems have achieved remarkable capabilities in content creation. However, these same capabilities increasingly pose risks to children, including facilitating child sexual abuse and exploitation (CSAE), particularly via the development of AI-generated child sexual abuse material (AIG-CSAM). This misuse represents a critical challenge at the intersection of AI safety, child protection, and technical system design—one that most existing AI safety frameworks inadequately address [1]–[4].
The scope and severity of this problem are substantial. The National Center for Missing and Exploited Children (NCMEC) received more than 440,000 reports of AI-generated material related to CSAE in the first half of 2025 alone [5], compared to 7,000 reports in 2023-2024, and the Internet Watch Foundation (IWF) observed a 400% increase in AIG-CSAM reports since 2024 [2], [6]. A recent study found that 13% of U.S. sexual extortion victims reported that the perpetrator used AI to create blackmail material [7], and a 2025 survey of Australian young people found that 41% of sexual extortion victims reported that the material used to blackmail them had been digitally manipulated [8]. In the U.S., 1 in 17 teenagers aged 13–17 report having been victimized by deepfake nude images [9]. AIG-CSAM has thus been recognized by the AI safety community as a critical area for intervention, with immediate, quantifiable marginal risk [10].
However, despite the known risks of AIG-CSAM, many existing techniques from AI safety and secure & privacy-preserving machine learning cannot be directly used for AIG-CSAM prevention, detection, and mitigation. Paradoxically, in part because child safety is an important and established area of concern, there are unique constraints and restrictions surrounding CSAM access/usage. These restrictions are ethical and necessary, but create complications for researchers and practitioners seeking to develop and assess AI safety solutions. This motivates our main position (below), which we explore in detail in this work:
Main Position Existing AI safety research makes assumptions that do not match the legal and ethical constraints of safety applications preventing child sexual abuse and exploitation. Bridging this gap by solving open sociotechnical problems will be critical to make AI safety solutions effective for child safety in practice.
| General Public | AI Developers/Providers | Reporting Hotlines | Law Enforcement | |
|---|---|---|---|---|
| Can legally access CSAM | No | No* | Yes | Yes |
| Can access hashes | No | Yes | Yes | Yes |
| Can train on CSAM | No | No* | Yes | Yes |
| Can generate CSAM | No | No | No | Yes |
We outline current industry consensus on solutions to address AIG-CSAM [11], pointing out gaps between assumptions made by state-of-the-art mitigations and practical limitations. We then highlight 15 open problems (9 in the body, 6 in Appendix 10) spanning model development (\(\S\)3), deployment (\(\S\)4), and maintenance (\(\S\)5), which represent critical research and policy priorities where advances could have immediate impact on protecting children. Finally, we provide recommendations for researchers, AI providers, and policymakers to address these open problems and bridge ecosystem gaps.
Generative AI enables non-experts to create realistic digital media at unprecedented scale. The AI safety community has raised concerns about resulting risks, such as bias propagation [12], disinformation [13], bioterrorism [14], job displacement [15], and cybercrime [16]. In this work, we specifically examine generative AI risks for child safety, focusing on photo-realistic child sexual abuse material (CSAM). AIG-CSAM presents critical harms to child safety. It complicates victim identification by expanding the volume of material that law enforcement must navigate to locate children in abuse scenarios [1], [4]. It enables re-victimization as bad actors can fine-tune models on existing CSAM to generate content imitating specific victims while producing new poses and acts of violence [1], [4]. It also broadens the pool of victims for sexual extortion, with offenders using image editing tools to sexualize benign depictions of children for such schemes [17].
Both preventive and reactive measures around AIG-CSAM are necessary to improve child safety: civil society [18], industry [19], and regulatory bodies [20] have all pursued relevant efforts. However, as we show, CSAM imposes unique constraints on existing AI safety techniques. To illustrate the resulting open research gaps, in this section we first describe what makes AI child safety particularly challenging (\(\S\)2.1) and detail key organizations involved in intervention efforts (\(\S\)2.2). Disclaimer: Throughout this work, we note that we focus on AI-generated imagery as the key modality of interest, and primarily take a U.S./western perspective when considering legal constraints that may impact safety techniques. We discuss limitations, broader perspectives, and alternate views in \(\S\) 6 and \(\S\) 7, and provide a detailed discussion of related work in App. 9.
The illegal and sensitive nature of CSAM, alongside other guardrails related to children’s data, creates unique challenges for AI safety. Below we highlight several underlying challenges which we will revisit throughout the paper:
— Data restrictions. CSAM is illegal in most countries, with legislation such as COPPA and GDPR establishing strict privacy safeguards for children’s data [21]–[23]. As we show, these data access restrictions are at odds with many standard AI safety approaches, which require access to training and/or evaluation data.
— Evaluation restrictions. Evaluation and red teaming methods typically rely on model prompting to assess model behavior and capabilities, but intentionally generating AIG-CSAM is also illegal in the U.S. [24]. Similarly, standardized CSAM evaluation datasets do not exist, as it is illegal to access this data. — Adversarial and opaque environment with strict guarantees. Child safety demands strong guarantees, as AIG-CSAM output is intrinsically illegal in most countries and can directly harm specific children. However, these guarantees can be difficult to achieve, as CSAE offenders collaborate to actively circumvent safeguards, while AI system developers typically only disclose high-level descriptions of safety mechanisms (\(\S\)10.2).
— Wellness implications. CSAM exposure and the psychological demands of adversarial assessment pose substantial risks, including trauma and PTSD [25]. This can make it challenging to apply safety techniques requiring significant human involvement.
Three key groups make up the AIG-CSAM prevention and response ecosystem (see Table 1 for a summary).
AI Developers and Providers. AI developers (individuals/organizations that build AI technology) and AI providers (platforms that host AI tools) both have CSAM reporting and retention obligations under U.S. law [26]. They can access CSAM hashes via institutions like NCMEC for matching and reporting. In certain cases, they may be allowed to extract embeddings from CSAM detected on their platforms to train detection methods [27]*, but direct access would require partnership with law enforcement or reporting hotlines.
Reporting Hotlines. Hotlines (e.g. NCMEC, IWF) are authorized to receive and process CSAM reports, to facilitate removal from online platforms and support LE investigation and prosecution efforts. These organizations legally house CSAM, enabling hash sharing, victim ID, and abuse prevention efforts. While they can provide scoped, secure access to CSAM to other institutions building CSAM prevention and detection technology, they cannot directly assess generative AI models for CSAM capabilities.
Law Enforcement (LE). Finally, LE has the legal authority to investigate, collect, and analyze CSAM as evidence while working to identify victims, apprehend perpetrators, and disrupt child exploitation networks. However, although they have the ability to train detection methods and assess AI models for CSAM capabilities, they may lack the budget or technical expertise for these efforts.
As we will see, the restrictions surrounding CSAM access for these various groups (particularly for AI developers/providers) have wide-ranging impacts on the use of existing safety tools related to AI system development (\(\S\)3), deployment (\(\S\)4), and maintenance (\(\S\)5).
Developing safe generative AI models requires proactively addressing child safety risks before and during training. Large-scale datasets used to train generative models are often scraped from the Internet with minimal cleaning [28], leading to cases where datasets used to train popular models were found to contain CSAM [29], [30]. Generative models may also produce harmful outputs through concept fusion, combining attributes from separate training examples (e.g., adult pornographic content and benign child depictions) [31], [32]. Furthermore, by using generative models to partially edit CSAM, offenders can make it difficult to determine whether material depicts active abuse, recovered victim-survivors, or deepfakes.
Currently, model developers can use allowed/disallowed website lists for data curation to avoid sites hosting CSAM, [33], and can detect CSAM in training data by hash matching against third-party CSAM databases and using CSAM classifiers [18], [27], [34]. One approach to address concept fusion is NSFW filtering, which considers removing adult material from datasets prior to model development [11].
Developers also use red teaming to find harmful content queries, patching exploits before model release [35], [36]. Directly prompting for AIG-CSAM is illegal in the U.S., so developers report either testing proxy concepts [11] or avoiding such testing entirely [37].
Content provenance solutions have also been implemented to support victim ID efforts. These tools provide a feedback mechanism for researchers, AI providers, and policy makers. By allowing stakeholders to become aware of the models that are producing problematic output, systemic issues can be addressed in existing mitigations. Current approaches include C2PA [38] (an open metadata-based standard for digital content origin and edits) and watermarking [39]–[41].
Open Problem A1: Partial data cleaning How does data cleaning affect a generative model’s ability to depict a concept? To what degree can partial cleaning guarantee that a model cannot depict CSAM?
Guaranteeing complete CSAM removal is difficult due to imperfect detection technology, scale of data, and moderator wellness concerns. Determining whether imperfect removal (e.g., 99.9%) can prevent text-to-image models from learning to generate CSAM is an important open question. However, direct analysis of CSAM filtering strategies is challenging due to data access restrictions.
Existing work. LLM [42], [43] and diffusion studies [44] show that filtering training data can minimize harmful capabilities in the model, and research in text-to-image generation shows a critical number of training samples are required for concept composition [31], [45].
Limitations. AIG-CSAM generation requires strong safety guarantees. While existing work is a starting point to study the effectiveness of data cleaning, the AIG-CSAM problem would benefit from formal guarantees on a model’s ability to generate harmful content. Moreover, because it is illegal for researchers to generate CSAM, it is not yet clear how to exhaustively test models’ capabilities to ensure they are safe in these scenarios.
Open Problem A2: Preventing concept fusion Can generative models be selectively blocked from combining high-risk concepts, such as children and NSFW material?
A unique concern for CSAM is that harmful content can also be created by composing two potentially benign concepts (e.g., benign child depictions and adult content) (see example experiments in \(\S\)3.2.1). Addressing this via comprehensive data filtering is challenging for similar reasons to CSAM removal. In addition understanding whether partial data cleaning prevents concept fusion, novel architectures and training paradigms could also be developed to selectively prevent unwanted fusion.
Existing work. Prior work includes retrieval-based architectures for revocable sensitive data access [46] and classifiers that self-identify concepts to generate each output [47], [48]. Recent work also explores concept-based explainability for text-to-image models via sparse autoencoders [49], [50].
Limitations. Most concept-based explainability work requires accessing and generating images to identify interpretable features. Mechanisms relying on isolated sensitive datastores face challenges in reliably classifying sensitive imagery (e.g. NSFW content), may make sacrifices in quality on benign concepts, and are further complicated by entangled concepts (e.g. children’s medical imagery).
A concern we raise in Open Problem 2 is that models may be able to combine high-risk concepts such as children and NSFW material. As shown in Figure 1, we find this to be the case with proxy concepts from CelebA [51]: we train diffusion models on separate images of people with blonde hair and eyeglasses, and find that they can generate the combined concept of blondes wearing eyeglasses with zero prior examples. Concurrent work from [45] similarly uses eyeglasses as a proxy for nudity, and finds that text-to-image models can generate children wearing eyeglasses even with 94% of child images removed. Developing alternative techniques to prevent concept fusion is thus an important area of future work. See App. 10.1.1 for full experiment details.
Open Problem A3: Resilience to harmful fine-tuning How can generative models be post-trained to prevent fine-tuning on CSAM or simultaneously fine-tuning on multiple CSAM-related concepts?
Unfortunately, even if base models appear ‘safe’, users can also unlock harmful capabilities through post-training procedures such as fine-tuning. Open-weight models allow users to adapt models on various tasks and are vulnerable to tampering [52]. CSAM perpetrators use GUI-based LoRA fine-tuning software like ComfyUI [53] or Ostris [54] to locally fine-tune open source models on CSAM [55]. “Nudifying” apps generating illegal deepfake sexual material of children similarly exploit the open source ecosystem, optimizing models for clothing removal and face-swapping [56]. Ideally, text-to-image models would allow benign fine-tuning while preventing these harmful uses.
Existing work. [57] propose self-destructing classifiers that degrade when fine-tuned on specific tasks. Similar approaches have been explored for diffusion models [58], [59].
Limitations. Solutions for self-destructing models that obstruct joint fine-tuning on separate concepts (e.g., adult sexual content and children) without blocking individual concept fine-tuning are largely unexplored. Similarly, strategies to prevent harmful LoRA fine-tuning (vs. full fine-tuning) are lacking. Methods that build resilience to harmful fine-tuning often require training on obstructed tasks; CSAM data access restrictions make this challenging.
We explore further problems in development, such as watermarking and minimizing human exposure, in App. 10.1.


Figure 1: Conditional diffusion models trained on images of blondes (top) and people in eyeglasses (middle) can generate blondes wearing eyeglasses (bottom), without overlapping training data. Fusion is effective with a critical threshold of training data (right)..
Below we outline take-aways for researchers, AI providers, and policymakers based on the identified open problems.
Research Directions
Assess partial data cleaning limits using proxy datasets; explore architectures preventing unwanted concept fusion for images/video.
Build self-destructing models targeting the “nudifying” ecosystem.
Create mechanisms to obstruct harmful fine-tuning that do not require access to the harmful material.
Develop robust content provenance: e.g. models that sharply degrade if finetuned to remove a watermark and localized provenance to handle inpainting.
Given data access restrictions, researchers and AI providers should also establish collaborations with reporting hotlines and LE to enable effective interventions and evaluation.
Concrete Steps for AI Providers
Engage CSAM survivors to understand their perspectives on partial data cleaning and re-victimization.
Partner with hotlines and LE for secure, scoped CSAM access to implement fine-tuning resilience.
Reliably label model outputs with C2PA or an equivalent standard and indelible watermarks. For open source models, prioritize tamper-proof content provenance solutions.
Policymakers play a critical role by establishing long-term pathways to encourage safer model designs. Regulatory efforts, research funding, and private-public partnerships that prioritize child safety are all important for effective action.
Policy Goals
Establish avenues for scoped, secure evaluation of AIG-CSAM capabilities by vetted institutions.
Resource standards organizations (e.g. NIST) to ensure that their guidance for content provenance in the open source setting remains up to date.
Instruct regulators to engage with the fine-tuning software ecosystem; assess what interventions are in place to prevent the misuse of these tools.
Task grant-making institutions to prioritize related research efforts in AI child safety.
Once AI models have been trained and evaluated for child safety, they should also be deployed with safeguards built into the release and distribution processes. Safeguards may include techniques such as content moderation, usage violation detection, reporting mechanisms, and transparency disclosures. These defenses are complicated by the AI model ecosystem, with providers building iteratively off each other’s systems via post-training and distillation.
Closed-source AI providers currently use safety classifiers [60] to detect harmful inputs/outputs. They can direct users who attempt to generate illegal content to help resources [33]. Some open- and closed-source providers establish user reporting pathways for violating outputs, prompts, and models. Open-source models may also include safety classifiers [61], or be safety fine-tuned to reject adversarial prompts [62], [63].
To anticipate adversaries who may attempt to circumvent safeguards, providers employ red teaming to identify harmful prompts for training classifiers [35]. However, this is less common on third-party platforms [33]. While platforms like HuggingFace generally prohibit third-party models with AIG-CSAM capabilities, enforcement of these policies varies [64]. One measure for developer transparency and accountability is recommending model cards.
Adversaries can jailbreak systems by disguising harmful signals [65], [66]; text-to-image models are particularly vulnerable [67]–[70].
Open Problem B1: Effective prompt/output detection How can harmful prompts/outputs be reliably detected, in an adversarial ecosystem with limited access to in-distribution examples?
Open source deployments face additional challenges: content moderation is not feasible, safety filters are vulnerable [71], and fine-tuning interventions are easily reversed [72], [73].
Existing work. Existing harmful prompt refusal requires instruction fine-tuning on similar prompts or few-shot examples in context [69], [74]. Other solutions include automated red teaming tools to search for violatory prompts [75], zero-shot prompt classifiers [60] and guideline based model training to reject unsafe prompts [76].
Limitations. Red teaming prompts and guidelines may not reflect actual CSAM offender behavior. Most organizations do not have access to CSAM; accurate classifiers require offender data. For those that do, there is no public benchmark to evaluate their solutions. Research to strengthen solutions in the open source setting is broadly lacking.
Open Problem B2: Automated model assessment How can models be assessed for AIG-CSAM capabilities and CSAM training data automatically?
Third-party models constitute a significant chunk of the model ecosystem, and are not always assessed for CSAM risks pre-deployment. Scalable assessment of deployed models is needed to enforce platform policies and prevent distribution of models with AIG-CSAM capabilities.
Existing work. Training data extraction techniques [77] have been applied to detect the presence of specific media in training data. Data attribution techniques for generative models could also prove useful to identify the samples that enable AIG-CSAM [78]–[80]. The nascent area of mechanistic interpretability (MI) aims to examine learned weights to understand large models [81], with techniques such as automated circuit discovery attempting to identify computational subgraphs that implement specific model behaviors [82], allowing auditing of specific capabilities by searching relevant subgraphs.
Limitations. Research using MI to audit CSAM generation capabilities is unexplored; seeking out CSAM trained models to assess such techniques may open researchers to risk [83]. Current techniques for training data extraction rely on prompting for AIG-CSAM, which violates US law. Text-to-video assessment is also broadly unexplored.
We explore further problems in deployment, such as model transparency and standardize assessments, in App. 10.2.
As with the previous open problems, in the absence of collaborations that enable researchers access to CSAM or CSAM trained models, there are still lines of research that can be pursued that directly ladder up, e.g.:
Research Directions
Design image-free auditing to enable upstream detection (e.g. auditing before diffusion denoising completion [84], [85].)
Harden models to white-box attacks using natural-language safety specifications [76].
Develop training data extraction techniques that do not rely on direct prompting or model training
Explore concept fusion via MI to identify and downweight neurons enabling adult-child combinations.
AI Providers have unique visibility into how offenders are misusing their platforms. They should prioritize efforts to use and share this data and other related metadata available to strengthen safeguards, including efforts such as:
AI Providers
Join data sharing programs (e.g. [86]) with trusted organizations to broaden access to adversarial prompts across platforms.
Deploy metadata and other signals to detect policy-violating models. Resource teams to meet the volume of models uploaded/downloaded.
Partner with external teams pre-deployment for model evaluation, sharing platform-specific offender behavior data for effective assessment
Policymakers are key to establishing broader systems and structures for building public trust in safety technologies used to safeguard generative models. They also play a critical role in drawing clear lines in the sand that allow for scalable, preventative efforts.
Policymakers
Resource institutions like NIST to establish pathways for public benchmarking of safety tech.
Pass legislation creating liability for intentional development or distribution of models built to produce AIG-CSAM or “nudify” pictures of children.
Require that developers disclose whether they conducted CSAM filtering on their training data (e.g. [87], [88])
Finally, model monitoring and maintenance is necessary to address emerging and evolving threats to children, maintain efficacy against an evolving technology stack, and reflect the broad nature of the ecosystem (where interweaving technology providers and systems may inadvertently reinforce harms, or create new harms [89]). In particular, open source models are highly vulnerable to abliteration: fine-tuning to remove safety guardrails [90], allowing offenders to enhance CSAM generation or “nudify” minors. Technology aggregators (model hosting platforms, app stores, search engines, AI character platforms, etc.) enable easy discovery of these harmful models [91], [92].
Current industry safety solutions for detecting harmful models are primarily reactive—removing flagged models rather than evaluating before providing access. While “nudifying” applications can be easily discovered via simple text matching searches [93], interventions to delist or suppress such sites and resources are largely lacking. Those that do incorporate detection strategies report using model hashlists (blocking uploads of files that are on hashlists of known CSAM optimized models [33], [94]) or similar efforts such as removing advertisements of “nudifying” applications [95]. Industry further reports efforts to maintain the quality of their own detection technologies, include child safety policies for each of their services, and collaborate with child safety organizations [33].
Open Problem C1: Identifying abliterated models How can we identify models and services that have been optimized for CSAM and “nudification”?
With thousands of models, apps, and services created and uploaded daily, rapid identification of those optimized for harmful purposes remains a critical gap.
Existing work. Model fingerprinting uses adversarial attacks to compare outputs between original and suspected stolen models for IP protection [96]. Model diffing exploits mechanistic differences between base and fine-tuned models to identify model-specific concepts [97], [98]. AIG-CSAM model hashlists can be built using cryptographic hashing. Limitations. Model fingerprinting requires insight into the “original” model, which is challenging to determine for models optimized by malicious actors. Model diffing is not well explored for text-to-image models, particularly for LoRA fine-tuning (often used by CSAM perpetrators). Cryptographic hashing detects exact model replicas but not minor modifications. Further, sourcing such hashlists can require accessing Tor onion sites dedicated to child abuse, which comes with significant legal and wellness risks.
Open Problem C2: Robust unlearning How can we reliably erase the concept of CSAM from generative models?
CSAM might be found in training data after a model is deployed; concept fusion opens up other pathways for CSAM generation. The gold standard is to retrain the model from scratch without harmful data, but retraining can be prohibitively expensive and time-consuming [99]. If the model was built by a different developer than the person who discovered the issue, engaging the original developer to conduct a full re-train may be challenging.
Limitations. Existing work has extensively explored approximate machine unlearning or concept erasure for text-to-image models, to remove knowledge with only a textual description of the harmful concept [63], [100]. However, these methods are vulnerable to adversarial prompts [101] and only provide probabilistic guarantees that a model no longer retains certain knowledge. CSAM requires exact unlearning, providing strict guarantees that CSAM images have no effect on model output [99], [102].
Open Problem C3: Protecting user imagery How do we proactively protect users’ imagery from unwanted AI-generated manipulation?
Solutions at the model level to prevent adversarial optimization are necessary, but do not afford end users (or platforms hosting user generated content) agency to proactively protect their own content from unwanted AI-manipulation. Existing work. Research exploring image immunization [103] involves injecting imperceptible perturbations into the image, such that image editing software fails to successfully edit the image. Some recent efforts specifically focus on protecting children’s imagery [104]. In the IP protection space, similar solutions have been proposed to disrput a model’s ability to mimic the style of particular artists [105].
Limitations. Some research indicates that these image perturbation strategies are not robust to simple attacks such as image upscaling [106]. Further, solutions to protect video imagery are lacking. Effectively evaluating these techniques ability to protect children’s imagery from “nudification” requires attempting to generate such imagery, which has ethical and legal implications.
Open Problem C4: Securing AI agents How can we prevent the misuse of AI agents and code generation to facilitate child sexual abuse?
Although criminal actors do not yet appear to be adopting automated code generation and AI agents for child sexual exploitation, the tools are already being used in other criminal enterprises, such as malware creation [107] and building hidden webcam recording software [108]. AI agents capable of relationship building could enable sexual extortion schemes; code generation could automate the creation of “nudifying” software, even if prebuilt nudifying apps are banned.
Existing work. Safety for agentic systems is an emerging field [109]. Current work focuses on early detection and prevention of misuse. As with traditional red teaming and filtering, most solutions rely on training models using examples of prior misuse, such as AI-generated malware [110]–[112]. Such examples are legal to develop and train on, and leading labs actively build this data.
Limitations. Developing “nudification” code or sexual extortion prompts is more ethically ambiguous than malware, raising data legality and researcher wellness concerns similar to directly red teaming for CSAM. Robust evidence that these emerging technologies are being misused in child safety contexts may be a prerequisite for industry prioritization, but obtaining the visibility needed to build such evidence remains challenging.
We discuss additional open problems in model maintenance, such as safeguard assessments, in App. 10.3.
Interventions for safe model maintenance are unique in that they can be both cross-platform and model specific, requiring research that covers broad surface areas for harm.
Research Directions
Build solutions to detect “nudifying” applications.
Develop model hashing techniques that are robust to minor modifications of the model.
Explore strategies that provide strong guarantees when remediating/unlearning harmful models.
Establish image/video protection techniques for chidren’s imagery, using proxy data and concepts.
AI Providers are already positioned to engage in both model specific and cross-platform safeguarding efforts:
AI Providers
Partner with cross-industry child safety organizations (e.g. the Tech Coalition) to enable cross-platform monitoring and measurement.
Partner with user-generated content platforms to evaluate strategies for protecting user imagery from unwanted AI-generated manipulation originating from your foundation models
Model hosting platforms jointly establish cross-industry consistency in policies and enforcement for identifying & removing harmful third-party models
Policymakers can create incentives and pathways such that stakeholders invest in cross-platform safety:
Policymakers
Task regulators to review “nudifying” services (websites, apps, etc.) for unfair, deceptive and fraudulent business practices that promote AIG-CSAM creation and distribution.
Mandate child safety specific impact and risk assessments for AI systems before deployment.
Direct regulators to assess platform (e.g. model hosting companies, search engines and app stores) strategies for preventing the distribution of models, apps and services that enable CSEA.
The most immediate alternative perspective is that child safety does not require fundamentally new AI safety approaches, but rather stronger or more consistently applied variants of existing approaches. Indeed, the core technical challenges (e.g., dataset governance, robust content filtering, adversarial testing, and post-deployment monitoring) are not unique to child safety and are broadly relevant to problems like disinformation, non-consensual imagery, and extremist content. Our work argues that despite the seeming similarities to these problems, the legal and ethical restrictions surrounding AIG-CSAM present non-trivial additional hurdles to existing approaches, necessitating in some cases entirely new techniques. While the problems identified are necessary to reduce the harms of AIG-CSAM, we believe many of them could prove beneficial for these other, less regulated domains.
Another view is that the problem of AIG-CSAM simply cannot be solved, or that we should instead focus AI safety efforts on other, existential risks [113]. Others may go further to suggest that AIG-CSAM is not harmful or is at least preferable to non-AI generated CSAM. Here, we note the significant evidence to the contrary [114], and advocate for a balanced position to invest efforts in multiple safety areas, especially as we could potentially make immediate headway on reducing the harms of AIG-CSAM during a time when reports are increasing sharply [2], [5], [6].
Finally, one may question whether technical AI safety interventions should be the primary locus of response at all. We could instead emphasize social, legal, and preventive measures such as digital literacy education for children and parents, victim support services, updated criminal law and international cooperation, and platform liability regimes that deter harmful deployment. From this standpoint, overemphasis on building new AI-specific safeguards may create a false sense of technological solvability while underinvesting in root causes. Others may go further to argue that policy solutions would eliminate the need for developing new technical approaches altogether, reducing the legal barriers to applying existing safety approaches. Our position reflects that policy solutions must come paired with, and be informed by, technical advancements.
In this work we argue that current safety techniques assume data accessibility, transparency, and evaluation practices that are in many cases incompatible with ethical and legal constraints surrounding CSAM. We then propose targeted recommendations and a set of 15 open technical problems that can improve child safety in the AI ecosystem.
We note limitations and areas for future work in our study. We focus on AI-generated imagery. However, exploitation spans multiple modalities, including text sexualizing minors [115], generative chatbots using celebrity voices for sexual conversations with children’s accounts [116], and video deepfakes of child sexual abuse [117]. While some of the issues identified in this work may also translate to other modalities, we leave the identification of technical open problems in non-image modalities as future work.
We also note that the risks and open problems identified throughout are primarily from a U.S.-centric perspective. The policy institutions, regulatory bodies, and frameworks referenced are also primarily U.S.-centric. Global regulation in this space is actively evolving [118], [119], and those evolutions may add additional nuance and context to the content in this work. A short discussion of policy tradeoffs is provided in Appendix 11, but exact policy implementations are left for future work and discussion with policymakers.
Moving forward, measuring and reducing the utility cost of existing safety methods is critical for adoption. At present, most AI-generated abusive material and models optimized for AIG-CSAM derive from open-source models. As such, the most pressing open problems are those that address the adversarial misuse of open-source models. Resilience to harmful fine-tuning (A3) is imperative, as perpetrators actively use GUI-based LoRA tools to abliterate models for AIG-CSAM generation. Other urgent problems include watermark robustness (A5), moderating the hobbyist ecosystem (B5), detecting abliterated models (C1), and robust unlearning (C2); we direct readers to [52] on broader open-source safety risks. Establishing collaborative venues (e.g., conference workshops) that foster communication between stakeholders could accelerate progress on these open problems.
We thank Steven Wu for his contributions in early planning for this work, and Stephen Casper for his feedback on later drafts. We also thank Tim O’Gorman, James Williams, Jim Pitkow, and Melissa Stroebel from Thorn for their reviews and comments on the paper.
| CSEA | Child sexual abuse and exploitation |
| AIG-CSAM | AI-generated child sexual abuse material |
| Reporting hotlines | National Center for Missing & Exploited Children (NCMEC), Internet Watch Foundation (IWF) |
| AI Developer | The entity responsible for designing, training, and testing an AI model or system. |
| AI Deployer | The entity that puts an AI system into service and controls its operation, often by integrating it into a product or platform. |
| AI Provider | Any entity that makes an AI system or model available for use. |
| Open Problem # | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Data restrictions | |||||||||||||||
| Evaluation restrictions | |||||||||||||||
| Adversarial environment | |||||||||||||||
| Wellness implications |
Extending AI safety to AIG-CSAM prevention naturally builds on diverse prior work that mitigates AI risks. Throughout this work, we highlight many relevant papers in these domains, e.g., on topics such as on red teaming [35], model fingerprinting [96], and tamper resistance [57]. In addition, several papers (1) outline broader open problems in AI safety, or (2) propose child-safety specific mitigations. We highlight these areas of related work here.
[52] outline open problems in derisking open-source models, several of which directly address child safety, such as reliable tamper resistance (Open Problem 3: Resilience to harmful fine-tuning). Similarly, [117] demonstrate problems with the open-source model landscape that enable video deepfake propagation, including video CSAM. Specific to Open Problem 1 (Partial data cleaning), [45] find that partial data filtering is insufficient for CSAM prevention and deconstruct open subproblems, such as accurate child detection.
Some existing works specifically address open problems outlined in this paper for preventing AIG-CSAM. For example, [30] propose annotation pipelines to improve data cleaning (Open Problem 1) and apply them to Stable Diffusion. Work from [120] develops an end-to-end CSAM classifier and proposes an evaluation methodology incorporating proxy data and controlled access to CSAM (Open Problem 6). There remains a significant need for further work on these problems.
Child sexual abuse material (CSAM) is unique in that it is the only media in the United States that is illegal to create, possess, or distribute [121]. This is also broadly true under global regulatory regimes [122]. While there are other domains with data that may be similarly traumatic to view, CSAM is unique in its illegal status. Government-sensitive material bears some similarity in that its distribution is limited [123], but CSAM access is not tied to security clearance levels the way other government-sensitive data is. Medical records under HIPAA also bear similarity in that data is restricted [124], but they are not subject to evaluation, generation, and wellness constraints that CSAM falls under.
Open Problem A4: Minimizing human exposure How does exposure to AIG-CSAM affect red teamers? How can human exposure be minimized?
The emotional and mental toll of CSAM exposure is well documented across moderation [125], law enforcement work [126], and data labeling [127], with additional wellness implications for red teamers [128]. These harms demonstrate the clear need for red teaming solutions that minimize human exposure to CSAM.
Existing work. Several frameworks for automated red teaming text-to-image models have been proposed [75], [129], [130], and existing industry content review tools incorporate wellness features like image blurring [131]. Limitations. Text-to-image models are particularly vulnerable to out-of-distribution prompts, requiring robust human red teaming [67], [132]. Offenders quickly discover new jailbreaking mechanisms, making it challenging to ensure testing remains relevant. Moreover, evaluating wellness features for red-teamers requires sociotechnical studies with industry and red teaming service providers, who may not have incentive to participate in such studies.
Open Problem A5: Watermark robustness How can text-to-image watermarks be made robust to removal, spoofing, and partial edits?
In the CSAM space, offenders actively combat safety interventions, including stripping metadata and other identifying factors in the image (e.g. watermarks). Researchers have explored robust watermarking techniques that are less susceptible to spoofing or removal [133], [134]. Methods like Stable Signature incorporate the watermark directly in the model weights during pretraining [41].
Limitations. Many schemes remain vulnerable to simple attacks [135]. Watermarks can be removed by partially modifying an image [136] or editing code in image generation scripts [137]. Watermarking techniques that are more robust to fine-tuning and erasure tend to be more vulnerable to being stolen and spoofed on other images [135]. Robust solutions that include a history of partial edits are also lacking.
To demonstrate concept fusion, we trained conditional flow matching models on 128x128 CelebA images. Each model comprised a 295M parameter UNet and was trained for 700 epochs. Given two attribute classes \(A, B\) (e.g., blonde hair, eyeglasses), models were trained on images from \(A \setminus B \cup B \setminus A\)—for instance, blondes without eyeglasses and non-blondes with eyeglasses (see Section 3.2.1 for examples). We train on 4K images of a certain hair color (blonde, black) without eyeglasses and a varying number of images (64-4K) without that hair color but with eyeglasses.
We measured the models’ propensity to generate compositional images from \(A \cap B\) (blonde hair and eyeglasses in Figure 1, black hair and eyeglasses in Figure 2) while varying dataset composition. In Figure 1, we test two generation methods of (1) unconditional generation or (2) conditioning on an average class vector. We then take the maximum detection rate over these two methods. We observe a sharp threshold in the number of samples required for composition: with 750 eyeglasses samples, composition likelihood was \({\sim}0.1\%\); with 1000 samples, this increased \(15\times\) to \(1.5\%\).
Varying the sampling strategy can further increase this ratio. In Figure 2, we evaluate a sequential conditioning approach: instead of denoising unconditionally for \(T\) steps, we first denoise for \(n < T\) steps conditioned on class \(A\), then for \(T - n\) steps conditioned on class \(B\). While a fixed conditioning (unconditional or average) measures propensity, this strategy measures capability. Here, capability proves substantially stronger—the model generates images from \(A \cap B\) over \(25\%\) of the time with just 250 eyeglasses samples (\(6.2\%\) of training data).
Our findings align with [31], who also observe a sharp sample threshold for concept composition. This suggests partial data cleaning may be effective, though the bar is high: in our setting, just 64 images of a “cleaned concept” still enable composition \({\sim}10\%\) of the time. Future work should explore how this threshold scales with larger diffusion models. Concurrent work from [45] more directly evaluates model capability to compose children with other concepts and highlights additional challenges such as child detection. We note both directions as ways to pursue child-safety research on the open problems posed here, at different levels of abstraction.
Open Problem B3: Standardized safety assessments How can we standardize assessments for AIG-CSAM capabilities?
Standardized safety assessments allow for consistent and transparent model evaluation. For child safety in particular, building confidence and trust in evaluations requires assurance that the assessment is robust, and not unduly influenced by other incentives, e.g., product deadlines.
Existing work. External red teaming and benchmarking are standard for assessing generative models. External domain expertise can help discover novel issues [138]. Benchmarking supports scalable reproducibility, allowing multiple models from different developers to undergo the same evaluation [139].
Limitations. Safety benchmarking assessments may correlate with model capabilities rather than actual safety [140]. Other studies highlight fundamental gaps in AI safety assessments, particularly for non-text modalities [141]. Given the adversarial nature of CSAM offenders and the wellness and legal barriers, static benchmarks quickly becomes outdated as offenders develop new strategies for generating AIG-CSAM.
Open Problem B4: Model transparency How can we use model cards to encourage transparency without inadvertently enabling offenders?
Model cards are the industry standard for disclosing information about models. For AIG-CSAM, documenting child safety interventions creates a natural pause point for developers to assess safeguards.
Existing work. Model cards are intended to provide fair assessment on a variety of human critical factors, e.g., bias [142]. Model card format can influence how interpretable the information is to non-technical audiences [143], enabling ethical decision-making for laypersons as well. Even the act of filling out model cards can elicit further ethical consideration from participating developers [144].
Limitations. Disclosing safety interventions is only effective if deployment is actually contingent on implementing them; currently, models without CSAM safeguards can still be released. While transparency enables good-faith actors to identify well-safeguarded models, it also enables offenders to discover vulnerable ones.
As a second proof-of-concept experiment, we manually audit public model and system cards for 14 widely used image and video generators (Table 4), asking whether they document child-safety relevant risks. These 14 models were chosen based on their presence in popular video generation [145] and image generation leader boards [146]. This analysis is limited to publicly available documentation and does not make assessments about any deployed safeguards, absence of documentation should not be interpreted as absence of safety measures. We mark whether a public card exists and whether it reports any evaluation targeted at children, explicit sexual content in general, and CSAM.
| Model | |||||
| Exists | Child | Explicit | CSAM | ||
| Sora | |||||
| Sora 2 | |||||
| DALLE-3 | |||||
| Imagen 4 | |||||
| Veo 3.1 | |||||
| FLUX.1 | |||||
| DeepSeek Janus | |||||
| SDXL | |||||
| Z-Image | |||||
| RunwayML Gen 4.5 | – | – | – | ||
| Grok Imagine | – | – | – | ||
| Midjourney | – | – | – | ||
| Pika | – | – | – | ||
| Luma | – | – | – |
Among the systems with model or system cards, only OpenAI’s Sora models include dedicated sections that explicitly discuss child safety and CSAM. Several of the most widely deployed systems (e.g., Midjourney, Grok Imagine, Pika, Luma) have no public model card at all, despite their scale and popularity.
Open Problem B5: Hobbyist ecosystem How can we encourage adoption of safety best practices across the diverse ML hobbyist ecosystem?
AI safety efforts typically target industry providers rather than hobbyists who build on foundation models [55]. This leaves a gap: downstream actors may lack the resources, incentives, or oversight to maintain the CSAM safeguards built into the original model.
Existing work. AI developers broadly recognize ethical dilemmas but often lack resources and training to navigate them [147]. Education-based prevention of child sexual abuse is well-studied and established as a standard approach [148], [149]. Research on hobbyist developers highlights intellectual stimulation as a primary motivation [150].
Limitations. Education efforts to promote developer awareness and CSAM prevention practices remain understudied. Researching those online communities carries risks of harassment and doxing [151], compounded by the high-stakes nature of child sexual abuse—a topic that often prompts defensiveness and avoidance when developers confront unintended consequences of their models.
Open Problem C5: Assessing safeguards How do third-party auditors and users effectively assess the efficacy of implemented safeguards?
Even where safeguards have been implemented or reported as implemented, assessing their effectiveness—individually and within the broader system context—is necessary for building trust and transparency.
Limitations. As noted earlier, most AI safety assessments focus on individual models through red teaming or benchmarks. Mechanisms for assessing complex AI systems are lacking. Sociotechnical assessments that account for how users actually engage with platforms, including offender behavior, platform-specific risks, and cross-platform dynamics, remain uncommon.
Companies may lack incentive to provide access necessary for these studies, which are important for building shared understanding and trust. Some metrics require cross-platform measurement, such as tracing circulating AIG-CSAM from downstream tools back to source models as a proxy for safeguard robustness.
Enacting the policy recommendations highlighted in this work may require difficult decisions and challenging tradeoffs. For example: establishing secure pathways for vetted institutions to evaluate AIG-CSAM capabilities could improve accountability and benchmarking, but would also require carefully scoped governance structures to prevent misuse, leakage, or other harms associated to exposure to highly sensitive material. When considering regulatory tools for transparency: mandatory disclosures on CSAM filtering, child-safety risk assessments, and platform prevention strategies all require clear standards for compliance auditing. In the absence of these, tradeoffs between flexibility and enforceability may result in company disclosures that are overly vague and lack meaningful detail. Even with such standards in place, challenges emerge around establishing the right level of specificity within these standards. Narrow requirements risk becoming quickly outdated as AI systems evolve, while overly broad requirements could create inconsistent enforcement.
These examples do not reflect the full scope of tradeoffs and decisions that may need to occur when enacting policy solutions in this space. Even so: we emphasize that, while these decisions are important and challenging, not enacting policy solutions is just as much a decision. The outcome of that decision, when it comes to AIG-CSAM, is tangibly apparent: more victims and more harm. This outcome is unacceptable.